Anthropic Ships Claude Haiku 5.5 at $0.10/$0.50, With a 5x Price Jump Past 100K Tokens
TL;DR
Anthropic released Claude Haiku 5.5 on October 7, 2026, almost a year after Haiku 4.5. The headline price is $0.10 per million input tokens and $0.50 per million output tokens, a 90% cut from Haiku 4.5's $1/$5 and an exact match for OpenAI's GPT-6 Luna. The catch is in the pricing table: that rate only applies to prompts up to 100,000 tokens. Past that line, input and output both cost five times as much, $0.50/$2.50. Haiku 5.5 also moves to the newer tokenizer that counts roughly 30% more tokens for the same text, which is why Anthropic's own average-savings claim is 75%, not 90%. On capability it is a different model: 39.2% on Terminal-Bench 4.0 where Haiku 4.5 scored 0.0% and Luna 16.4%, and 72.4% on OSWorld 2.1 against Luna's 48.9%. Every one of those numbers was run at max effort, and the API default is medium.
The price has a cliff at 100,000 tokens
Anthropic's pricing page lists Haiku 5.5 as two rows, and the second row is the one that matters for anyone running agents. For prompts up to 100,000 tokens: $0.10 input, $0.50 output, $0.125 for a five-minute cache write, $0.01 for a cache read. For prompts over 100,000 tokens: $0.50 input, $2.50 output, $0.625 cache write, $0.05 cache read. Batch is half price in both tiers. The docs say this out loud in the long-context section: every other Claude 4.6-or-later model bills its full 1M window at one rate, and Haiku 5.5 is the single exception.
Anthropic's justification is a usage statistic: prompts up to 100,000 tokens "make up around 90% of requests to our previous Haiku model." That is probably true for the classification, extraction, and routing jobs Haiku is marketed for. It is much less true for the other job Anthropic names on the model page, subagent work, because an agent's prompt is its accumulated context. A coding session that starts at 20K tokens and crawls past 100K does not just pay more for the tokens past the line. Every subsequent turn, output included, bills at the high tier.
Think of a parking garage that charges a dollar an hour for the first two hours and then re-rates the whole stay at the full-day price, and your agent never leaves the garage. The cache-read rates are tiered the same way, which reads as the whole prompt, cached prefix included, deciding the rate.
For comparison, OpenAI's model page lists Luna at a flat $0.10 input, $0.50 output, and $0.01 cached input across a 1,050,000-token context. Below 100K the two are priced identically. Above it, Luna is the cheaper small model by a wide margin, so long-context routing between the two is now a pricing decision as much as a quality one.
The tokenizer eats part of the discount
Haiku 5.5 uses the tokenizer introduced with Claude Opus 4.7, and the What's new page is blunt about it: "the same input text produces approximately 30% more tokens on Claude Haiku 5.5 than on Claude Haiku 4.5." Simon Willison measured around 1.25x on his own prompts. This is why the announcement says Haiku 5.5 costs "75% less to run" on average rather than 90%: the 75% figure folds the tokenizer penalty into the price cut.
Our arithmetic, using Anthropic's approximate 30% figure, for a request with a 50,000-token prompt (counted on Haiku 4.5's tokenizer) and 2,000 output tokens: Haiku 4.5 bills $0.060. Haiku 5.5 bills about $0.0078, roughly 87% less. Now make the prompt 150,000 tokens: Haiku 4.5 bills $0.160, Haiku 5.5 about $0.104, roughly 35% less, because the request lands in the top tier and the tokenizer inflation is charged at $0.50 instead of $0.10. The 90% headline is real for short prompts and shrinks fast once you cross the line.
Benchmarks: Luna loses, Sonnet still wins, everything at max effort
The announcement table compares Haiku 5.5 to Haiku 4.5, Sonnet 5.5, and GPT-6 Luna. Haiku 5.5 beats Luna on every row where Luna has a number, and loses to Sonnet 5.5 on every row.
- Terminal-Bench 4.0: Haiku 5.5 39.2%, Haiku 4.5 0.0%, Luna 16.4%, Sonnet 5.5 70.6%.
- OSWorld 2.1 (offline subset): 72.4%, 15.7%, 48.9%, 83.9%.
- GDPval-AA v2.1 (Elo): 1620, 735, 1437, 1840.
- AA-Briefcase v1.1 (Elo): 1578, 614, 1336, 1824.
- Humanity's Last Exam (no tools): 45.9%, 10.2%, Luna not reported, 56.9%.
- FrontierCode 1.1 (Main): 46.4%, no Haiku 4.5 result, 42.4%, 52.1%.
- Chartography (no tools): 46.4%, 6.4%, 29.1%, 61.6%.
The system card is where the fine print lives. All Haiku 5.5 results use "adaptive thinking at max effort, default sampling settings (temperature, top_p), averaged over five trials." Terminal-Bench 4.0 is 66 tasks, run ten times each for 660 trials, with a standard error of plus or minus 1.9 points. Safeguards stopped 12 of those 660 trials, ten of them on a single task, and every stopped trial counted as a failure. The eval ran with no internet egress. Luna's 16.4% comes from the public leaderboard with the Codex CLI at max effort, because, as far as Anthropic knows, OpenAI has not published its own Terminal-Bench 4.0 number for Luna. Haiku 4.5 passed none of its trials, which is at least consistent.
Two things follow. First, the API ships Haiku 5.5 with a default effort of medium, so the model you get without touching the effort parameter is not the model in the table, and the card does not publish a medium-effort Terminal-Bench score. Second, effort reorders the lineup in odd ways: on FrontierCode Main at max effort, Haiku 5.5's 46.4% edges Sonnet 5.5's 46.2%, and Sonnet only pulls ahead to 52.1% at xhigh. For a cheap model, max effort is where the bargains are, and max effort is also where the output tokens pile up.
What breaks when you swap the model ID
The model ID is claude-haiku-5-5 on the Claude API, Google Cloud, Microsoft Foundry, and Claude Platform on AWS, and anthropic.claude-haiku-5-5 on Bedrock. The context window goes from 200K to 1M tokens and max output from 64K to 128K (300K on the Batch API with a beta header). The reliable knowledge cutoff moves from February 2025 to June 2026. Anthropic commits to keeping it available until at least October 7, 2027. The What's new page lists the breaking changes, and there are more of them than a point release usually carries:
- Non-default
temperature,top_p, ortop_kreturns a 400 error. Omit them. - Manual extended thinking with
budget_tokensreturns an error. Adaptive thinking is on by default; steer it witheffort, or turn it off withthinking: {"type": "disabled"}at high effort or below. - Assistant message prefill returns an error. Messages must end with a user turn.
- Responses can begin with a
thinkingblock even when you never asked for thinking, so select content blocks bytype, not position. Thinking text is omitted by default; setthinking.displayto"summarized"to see it. - Thinking tokens count toward
max_tokens, so a small limit can stop after the thinking block and before any text. - Computer use needs
computer_toolset_20260801; the oldcomputer_20250124tool is rejected. - Safety classifiers can now return
stop_reason: "refusal", with no server-side fallback. - Editing earlier turns invalidates thinking blocks you send back, and thinking blocks only replay through the account that produced them.
One more date worth knowing: Anthropic's retirement commitment for Haiku 4.5 only runs to October 15, 2026. That is not an announced shutdown, but it is the point after which Anthropic is free to schedule one, and the migration list above is not a ten-minute job.
Sonnet 5.5 cache reads halved, and API credits for subscribers
Two side announcements ride along. Cache reads on Sonnet 5.5 drop from $0.20 to $0.10 per million tokens, which Anthropic says "reduces the cost of Sonnet 5.5 on most agentic tasks by around 20%." The pricing page confirms the 0.05x multiplier; Opus 5.5 keeps its $0.20 cache read. If you run long Sonnet agent loops, this change is worth more than Haiku is.
Claude subscribers also get monthly API credits on the Claude Platform: $100 for Max 5x, $200 for Max 20x, and up to $500 pooled for Team subscribers. Willison notes the credits do not roll over. At Haiku 5.5's low tier, $100 is a billion input tokens, so the practical effect is that Max subscribers can run a side project's classification and extraction pipeline for free.
Safety and the fine print
The system card is deliberately short; Anthropic says non-frontier models get condensed cards from here on. Its Responsible Scaling Policy section finds Haiku 5.5 "broadly less capable than Claude Opus 5 across domains," determines it does not cross the CB-2 or Autonomy-2 thresholds, and treats it as meeting CB-1 and Autonomy-1 with the matching mitigations. Misalignment risk is rated low, with a reason that doubles as a capability note: the model has "difficulty controlling its chain of thought or evading monitors" when its reasoning is visible. In cyber evaluations it "significantly outperformed" Haiku 4.5 but fell short of Opus 5.5, Mythos 5.1, and Opus 5.
The customer quotes in the announcement are the usual launch-day set, but a few carry numbers. Asana reports "over a 30% reduction in latency for task completions and up to 2.5x faster inference per agent turn." Box says Haiku 5.5 "scored 11 points higher than Haiku 4.5 at about half the latency." HubSpot reports 92.8% on an internal CRM suite averaged over three runs. Cognition says its Fusion system, with Haiku 5.5 as the sidekick model, holds a FrontierCode score of 66.2, which is the clearest hint at where a cheap, fast model goes in an agent stack: not driving, but doing the thousand small lookups the driver delegates.
Key Takeaways
- Haiku 5.5 costs $0.10/$0.50 per million tokens for prompts up to 100,000 tokens and $0.50/$2.50 above that. No other current Claude model has a long-context tier.
- The new tokenizer counts about 30% more tokens than Haiku 4.5, so Anthropic's average saving is 75%, not 90%, and the gap narrows further for long prompts.
- Agent sessions cross the 100K line as context accumulates, and every later turn then bills at the high tier. GPT-6 Luna is flat at $0.10/$0.50 across its whole context.
- Haiku 5.5 beats Luna on every published row (39.2% vs 16.4% on Terminal-Bench 4.0, 72.4% vs 48.9% on OSWorld 2.1) and trails Sonnet 5.5 on all of them.
- Every benchmark was run at max effort. The API default is medium, and there is no published medium-effort Terminal-Bench score.
- Migration is not a model-ID swap: sampling parameters, prefill, and budget_tokens all now error, and thinking blocks can lead a response.
- Sonnet 5.5 cache reads halve to $0.10, worth around 20% on agentic workloads, and Max subscribers get $100 to $200 a month in API credits.
Sources: Anthropic: Introducing Claude Haiku 5.5, Claude Haiku 5.5 System Card, Anthropic Pricing, Claude Haiku 5.5 model page, What's new in Claude Haiku 5.5, Claude Haiku 4.5 model page, OpenAI API: GPT-6 Luna, Simon Willison: Claude Haiku 5.5, VentureBeat