Claude Sonnet 5.5 Hits 70.6% on Terminal-Bench 4.0 at Sonnet 5's $2/$10 Price
TL;DR
Anthropic released Claude Sonnet 5.5 today, the second model in the Claude 5.5 family after Claude Opus 5.5. The price does not move: $2 per million input tokens and $10 per million output, same as Sonnet 5. The scores do: 70.6% on Terminal-Bench 4.0 (Sonnet 5 managed 10.3%), 55.5% on CursorBench 4.0, and a GDPval-AA Elo two points behind Opus 5.5, which costs twice as much per token. Anthropic also says it generates output 30%+ faster and costs up to 30% less per task than its predecessor. If your code runs Sonnet with thinking disabled, read the migration notes before you swap the model string, because disabled now returns a 400.
What shipped
The model ID is claude-sonnet-5-5, live today on the Claude Platform, the Claude apps, Claude Code, Amazon Bedrock, Google Cloud, and Microsoft Foundry. Per Anthropic's models overview, it gets the same 1M-token context window and 128K max output as Opus 5.5, with a June 2026 knowledge cutoff. Zero data retention is available, as with Opus 5.5 and Sonnet 5.
Anthropic positions it as the fast, cheap complement to Opus: strongest at "well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets," while Opus 5.5 "remains clearly stronger at complex, open-ended work requiring sustained judgment." Claude Haiku 5.5 is promised "in the coming weeks."
- Price: $2 input, $10 output, $0.20 cache reads, $2.50 cache writes per million tokens. Opus 5.5 is exactly double on everything except cache reads, which are $0.20 on both.
- Speed: "30%+ faster" output generation than Sonnet 5, which Anthropic calls its fastest Sonnet to date.
- Default effort: Medium in Claude Code and the apps, High on the Claude Platform.
The benchmark jump
The headline number is Terminal-Bench 4.0, where Sonnet 5.5 scores 70.6% against Sonnet 5's 10.3%. That is not a typo, and it is not a gentle generational bump either. It also edges out Opus 5.5's best Terminal-Bench score (66.4%, at Xhigh effort), which is the kind of result that should make you rerun your own evals before you trust your routing table.
The rest of Anthropic's table, all vendor-reported:
- FrontierCode 1.1 (main): 52.1% at Xhigh, 46.2% at Max, versus Sonnet 5 at 42.4%, Opus 5.5 at 54.4%, and GPT-6 Sol at 49.3%.
- GDPval-AA v2.1: 1844 Elo, versus 1846 for Opus 5.5, 1449 for Sonnet 5, and 1487 for GPT-6 Sol.
- AA-Briefcase v1.1: 1811, versus 1822 for Opus 5.5 and 1359 for Sonnet 5.
- Humanity's Last Exam (with tools): 64.5%, versus 67.7% for Opus 5.5 and 54.9% for Sonnet 5.
- Chartography (no tools): 61.6%, up from 15.6% on Sonnet 5, against 64.4% for Opus 5.5.
Two footnotes are worth reading. First, Sonnet 5.5 scores lower on FrontierCode at Max effort than at Xhigh. Anthropic says that at Max it more often ran Claude Code's code-review skill, which fans a review out across many subagents, and in cases Cognition examined that caused timeouts or out-of-scope edits the benchmark penalizes. More effort is not a free dial. Second, OSWorld is reported as a partial run, and the GDPval-AA and AA-Briefcase numbers came from a pre-release deployment with a since-fixed structured-outputs bug that Anthropic expects to understate, not inflate, the scores.
The real story is tokens per task
The list price did not change, so the savings have to come from somewhere, and Anthropic says they come from the model simply doing less. It "typically needs far fewer tokens to do the same work," batches tool calls more often than Sonnet 5, and takes fewer steps. On several benchmarks, Anthropic says Sonnet 5.5 at Low or Medium effort beats Sonnet 5's best score for roughly a tenth of the cost per task. At High effort on FrontierCode, it scores 10 points above Sonnet 5 at the same setting for about one fifteenth of the cost per task.
Think of it as the same hourly rate for a contractor who stopped rewriting the same function three times before showing you. The invoice shrinks without anyone touching the rate card.
The launch customers put numbers on it. Balyasny Asset Management ran 2,441 private finance tasks and reported about 121k tokens per answer for Sonnet 5.5 against 497k for Sonnet 5, with a higher score. Slack said its offline Slackbot evals improved on almost every measure with about 14% fewer output tokens and no prompt changes. Lovable reported a third fewer tool calls and roughly half the shell runs per task.
Base44 reported that across 118 real app builds, Sonnet 5.5 produced apps that scored level with Opus 5, averaging 3.6 iterations per build where Opus 5 took 7.7. Box said it was more accurate than the previous model, 2.4x faster, and used 12% fewer total tokens. Zendesk said tickets were processed 20% faster. These are customer quotes on a launch page, so treat them as directional, but they point the same way as Anthropic's cost curves.
Where it sits in the lineup
The price ladder is now clean: Sonnet 5.5 at $2/$10, Opus 5.5 at $4/$20, and Claude Fable 5.1 at $10/$50. The awkward part for Opus is that on CursorBench, GDPval-AA, OSWorld, and Chartography, the half-price model is within about three points.
Anthropic's own advice is that Sonnet 5.5 complements Opus 5.5 best at lower effort settings, where it is much cheaper per task, while at higher settings the two "can perform comparably at a similar cost." Translation: the cheap tier is cheap because it thinks less. Crank it to Max and you are paying Opus money for a model that, per the FrontierCode footnote, may start freelancing.
A sensible pattern, which Creator's Kevin Ngo describes in the launch quotes, is to let Opus set the architecture and hand the implementation to Sonnet. If you already run a planner/worker split, this is the model to try in the worker slot first.
Migration gotchas
The migration guide lists the breaking changes, and one of them will bite anyone who ran Sonnet with thinking off:
- Thinking is on by default. A request with no
thinkingfield runs adaptive thinking. Code that readscontent[0].textbreaks when the response opens with a thinking block, so read content blocks by type. disabledis gone. Sendingthinking: {"type": "disabled"}returns a 400. The replacement is{"type": "between_tools"}, the lowest setting, which skips up-front thinking but still emits short thinking blocks between tool calls.- Other 400s: thinking budgets, sampling parameters, assistant prefill, and forced tool choice are also rejected.
- Budget
max_tokensagain. It covers thinking plus text, and thinking tokens bill as output.
Safety and the new guardrails
Anthropic says Sonnet 5.5's cyber capabilities are comparable to Opus 5's, so it is the first Sonnet to ship with Opus-style cyber safeguards: routine bug-finding and fixing works, but higher-risk security tasks "will visibly fall back to Sonnet 5." Biology safeguards are unchanged from Sonnet 5, with the Life Sciences Verification Program for teams that need broader access.
It is also the first Sonnet with safety classifiers aimed at distillation, meaning attempts to extract its reasoning at scale through fake accounts. It extends preserved thinking so thinking blocks cannot be decoupled from the account that created them. Most developers will not notice, but if you move conversations between accounts, including switching accounts mid-session in Claude Code, check the docs. On the alignment side, the system card covers an automated audit of roughly 1,850 scenarios where Sonnet 5.5 matches or improves on Sonnet 5 on most measures, with Opus 5.5 still slightly ahead overall.
Key Takeaways
- Claude Sonnet 5.5 (
claude-sonnet-5-5) is live on the API, apps, Bedrock, Google Cloud, and Microsoft Foundry at Sonnet 5's unchanged $2/$10 per million tokens, with a 1M context window. - Vendor-reported scores: 70.6% on Terminal-Bench 4.0 (Sonnet 5: 10.3%), 55.5% on CursorBench 4.0, and 1844 on GDPval-AA, two points under Opus 5.5.
- The savings come from fewer tokens and steps per task: up to 30% cheaper per task and 30%+ faster output than Sonnet 5, per Anthropic.
- Max effort is not always best: Sonnet 5.5 scores lower on FrontierCode at Max than at Xhigh.
- Migration:
thinking: disablednow returns a 400; usebetween_toolsand read response blocks by type. - Haiku 5.5 is next, "in the coming weeks."
Sources: Anthropic: Introducing Claude Sonnet 5.5, Claude models overview, Sonnet 5.5 migration guide, Sonnet 5.5 system card, SiliconANGLE, The Decoder