Skip to content
chornous.dev

Type two or more letters. Esc closes.

Index of sheets
Theme

Sheet 03 · Writing

All notes

Claude Haiku 5.5: a tenth of Haiku 4.5's price under 100K tokens, and a tokenizer that counts 30% more

Anthropic released Claude Haiku 5.5 on October 7, 2026 at $0.10 input and $0.50 output per million tokens for prompts up to 100K. It adds effort levels, rejects thinking budgets, sampling parameters and prefill, and counts about 30% more tokens than Haiku 4.5 for the same text.

Pl. 89 · network drawing generated from the slug “claude-haiku-5-5-migration”

Anthropic released Claude Haiku 5.5 on October 7, 2026, nine days after Sonnet 5.5. The company calls it "the cheapest, fastest, and most capable small model we've ever released." The API ID is claude-haiku-5-5, a fixed ID with no date suffix and no separate alias. It runs on the Claude Platform, Amazon Bedrock, Google Cloud and Microsoft Foundry.

The launch post sells it for summaries, classification, compaction and subagent work under Opus 5.5 or Sonnet 5.5. If you run Haiku 4.5 in a queue worker today, the price cut is the headline. The migration guide holds the part that breaks your code.

Two price tiers

Haiku 5.5 splits its price at 100,000 prompt tokens. Haiku 4.5 had one price.

Price per million tokens, prompts up to 100K / over 100K
ItemHaiku 5.5Haiku 4.5
Input$0.10 / $0.50$1.00
Output$0.50 / $2.50$5.00
Cache writes$0.125 / $0.625$1.25
Cache reads$0.01 / $0.05$0.10

Anthropic puts the cut at 90% for requests up to 100K tokens and 50% above that. It also says about 90% of Haiku 4.5 requests fell under the threshold. The same announcement halved Sonnet 5.5's cache-read price to $0.10.

The tokenizer eats into those savings. Haiku 5.5 uses the tokenizer from Claude 4.7 and later, and the guide says the same text produces about 30% more tokens than on Haiku 4.5. Images cost more too: Haiku 5.5 uses the high-resolution tier, so a 2,000 by 1,500 pixel image costs about 2.5 times the visual tokens it did on 4.5.

A rough sketch of my own, for a 20,000-token prompt with a 1,000-token answer and no thinking: Haiku 4.5 charges $0.025. With 30% more tokens on both sides, Haiku 5.5 charges about $0.0033. Thinking tokens bill as output, so your number lands higher. Measure it.

The numbers

Terminal-Bench 4.0, Haiku 4.5 scored 0.0%
39.2%
OSWorld 2.1 offline subset, up from 15.7%
72.4%
more tokens for the same text
~30%

Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, and Anthropic still points you to Sonnet and Opus for complex agentic coding. Haiku 5.5 is also the first Haiku with effort levels, from Low through Max.

Five changes that return a 400

The migration guide lists eleven items for Haiku 4.5 users. These fail outright:

Haiku 4.5 request settings that Haiku 5.5 rejects
Haiku 4.5 settingHaiku 5.5 replacement
thinking type "enabled" with budget_tokensthinking type "adaptive" plus output_config.effort
temperature other than 1, top_p other than 0.99, any top_kremove them and steer with the prompt
a final assistant turn (prefill)end with a user turn, use structured outputs for format
computer_20250124 toolcomputer_toolset_20260801 on the Claude API and Google Cloud
edited system, tools or earlier messages with thinking blocks sent backkeep the conversation append-only

Sending temperature and top_p together also fails, even at their defaults. The append-only check applies in full to accounts created after August 31, 2026. Older accounts get the error after they opt in through thinking.block_binding.prefix_mismatch_behavior.

The thinking change is the one most Haiku code needs:

 {
-  "model": "claude-haiku-4-5",
+  "model": "claude-haiku-5-5",
   "max_tokens": 16000,
-  "thinking": { "type": "enabled", "budget_tokens": 8000 },
+  "thinking": { "type": "adaptive" },
+  "output_config": { "effort": "low" },
   "messages": [{ "role": "user", "content": "..." }]
 }

Behavior that changes without an error

Adaptive thinking is on by default. A request with no thinking field can come back with thinking blocks first, so code that reads content[0].text breaks. Those blocks arrive with an empty thinking field and a signature. Set "display": "summarized" if you log them.

Thinking tokens count toward max_tokens. A classifier tuned with max_tokens: 50 can stop with stop_reason: "max_tokens" before it writes any text. Raise the limit or drop to Low effort.

Thinking blocks from earlier turns now stay in context as input tokens. Haiku 4.5 dropped all but the latest turn's. Long chats grow faster than the tokenizer alone explains.

Forced tool_choice still works on Haiku 5.5, unlike on Sonnet 5.5. The response starts with the tool call and skips thinking. If you want the model to reason first, switch to auto and tell it in the prompt when to call the tool.

Haiku 5.5 runs safety classifiers and can return stop_reason: "refusal" with no fallback model. Priority Tier doesn't cover it, and Bedrock doesn't offer structured outputs for it.

This week

  1. Grep for budget_tokens, temperature, top_p and requests that end on an assistant turn. Those four fail first.
  2. Run your prompts through token counting with model: "claude-haiku-5-5" and recompute cost from those counts, not Haiku 4.5's.
  3. Read responses by block type, and handle refusal as a stop reason.
  4. Start your effort sweep at Low for classification and summaries, and raise it where your evals drop.

If you run Haiku 4.5 on Priority Tier and depend on that capacity, stay put until you've planned around it. In Claude Code, /claude-api migrate this project to claude-haiku-5-5 applies the mechanical changes and hands you a checklist for the rest.

Volodymyr Chornous