Claude Sonnet 5.5: Cheap Below Max, Not at the Top
Claude Sonnet 5.5, released on September 28, 2026, is the second model in the Claude 5.5 family. Anthropic does not sell it as a new frontier. It sells it as the faster complement to Opus 5.5 for well-scoped work: bug fixes, documents, slides, and spreadsheets. The list price did not change. $2 per million input tokens and $10 per million output, the same sticker as Sonnet 5 and as GPT-6 Sol.
The score that will get quoted is Terminal-Bench 4.0 at 70.6%, above Opus 5.5's best of 66.4%. That 70.6% is max effort, and Anthropic's own cost chart prices it at $12.54 a task. Opus reaches 66.4% at xhigh for $7.35. At medium, the default in Claude Code and the Claude apps, Sonnet scores 28.8% for $0.83 and Opus scores 57.6% for $2.94. The model is the cheap option below max. At max, it can be the expensive one.
Why this matters
On the Claude API the id is claude-sonnet-5-5. Context is 1 million tokens, max output is 128,000, and the reliable knowledge cutoff is June 2026. Claude Code and the Claude apps default to medium effort. The Claude Platform defaults to high. Those are different purchases. Haiku 5.5 is not this release.
A cache read is $0.20 on Sonnet 5.5 and $0.20 on Opus 5.5. Five-minute cache writes are $2.50 against $5. If a run is mostly cache hits, Sonnet is not half of Opus. The half-price applies to uncached input and output. Anthropic's claim of up to 30% less cost per task than Sonnet 5, and output more than 30% faster, is about fewer tokens on the same work. The tokenizer did not change.
Artificial Analysis prices a blended million tokens at $1.54, using seven cache hits, two uncached inputs, and one output. Their cost for one max-effort Intelligence Index task, with the default fallback left on, is $7.60. That index run produced 410 million tokens, against a median of 88 million. Anthropic's "fewer tokens" and that 410 million are different measurements. The index score there is 56, against 58 for Opus 5.5 at the same max-with-fallback setting. Near Opus on the composite, not past it.
What each effort actually buys
These points are Anthropic's accuracy-versus-cost charts, not the launch table. The table prints one cell. The chart prints five, and the cell is often the expensive end.
Terminal-Bench 4.0. Opus's 66.4% is xhigh, which Anthropic notes is that model's highest score on this bench. GPT-6 Sol was not public, so the chart uses GPT-5.6 Sol.
| Effort | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
| Medium | 28.8% at $0.83 | 57.6% at $2.94 |
| High | 43.0% at $1.94 | 64.2% at $3.88 |
| Xhigh | 61.5% at $5.30 | 66.4% at $7.35 |
| Max | 70.6% at $12.54 | 64.8% at $11.24 |
Medium Sonnet does beat Sonnet 5's best, 10.3% at $11.62, for less than a tenth of the cost. It does not beat Opus at any shared setting except max, and max costs more. Do not move terminal work off Opus because of the 70.6%.
FrontierCode 1.1, main set. This is the chart that should stop a max-effort default. High is the Platform default.
| Effort | Sonnet 5.5 | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|
| High | 49.4% at $0.42 | 54.0% at $1.09 | 47.7% at $1.04 |
| Xhigh | 52.1% at $1.59 | 51.4% at $2.25 | 48.4% at $1.32 |
| Max | 46.2% at $20.78 | 54.4% at $6.19 | 49.3% at $2.07 |
At high, Sonnet matches Sol's best, 49.3% at $2.07, for about a fifth of the cost, and it is 10 points above Sonnet 5 at the same setting (39.4% at $6.10). Opus at medium is already 54.6% at $0.80, above Sonnet's best. Sonnet at max scores 46.2% and costs $20.78, about thirteen times the xhigh bill, because FrontierCode penalizes edits outside the task. At max, Sonnet more often ran Claude Code's review skill across many subagents. In two cases Cognition looked at, that produced a timeout or extra edits, and the score fell.
CursorBench 4.0, real Cursor sessions. Sonnet at low effort is 35.8% at $0.50, above Sonnet 5's best of 34.1% at $7.17. Sonnet's best is 55.5% at $9.67. Opus at high is already 56.0% at $3.97, and Opus at max is 57.8% at $13.43. "Within about two points of Opus" is Sonnet's max against Opus's max. Opus gets a higher score at high for less than half that bill. GPT-6 Sol is not on this chart. GPT-5.6 Sol's max is 41.7% at $8.23.
AA-Briefcase v1.1, long-horizon knowledge work. Sonnet at medium is 1461 Elo at $1.64, above Sonnet 5's best of 1359 at $14.43, about a ninth of the cost. Sonnet at high is 1634 at $3.95, next to Opus at medium at 1642 and $4.40. Sonnet at max is 1811 at $29.19. Opus at max is 1822 at $21.05. At the top of this bench, Opus is slightly higher and cheaper, because Sonnet's max run spends more. GPT-6 Sol's max is 1483 at $2.67. Same $2 and $10 sticker as Sonnet, much lower Elo, much smaller bill at max.
GDPval-AA v2.1 stays the clean knowledge-work comparison, and it is not on these four charts. The launch table has Sonnet at 1844, Opus at 1846, Sonnet 5 at 1449, and GPT-6 Sol at 1487. Two Elo points under Opus, about 400 over Sonnet 5. Artificial Analysis ran that pair on a pre-release Sonnet that had a structured-output bug. Anthropic expects any effect to be small, and to understate the model. The bug is fixed.
The rest of the launch table, without a cost curve: Humanity's Last Exam with tools is 64.5% against Opus at 67.7%. OSWorld 2.1 is partial, 80.1% against 81.8%. Chartography without tools is 61.6% against 64.4%. BenchLM also lists a with-tools Chartography score of 90.2%, a separate setup, and an Artificial Analysis Terminal-Bench 4.0 of 63.6%. That 63.6% is not a correction of the 70.6%. BenchLM has 61 sourced rows and no overall score yet.

Anthropic's launch still of a one-file demo. The token count on the frame is part of the demo, not a benchmark.
What the testers are actually moving
Company tests, not independent benches. CodeRabbit says the over-search and the high token use of Sonnet 5 are gone, and plans to move simple and moderate reviews now. Balyasny, on 2,441 private finance tasks, puts Sonnet 5.5 ahead of Sonnet 5 at about 121,000 tokens an answer against 497,000. Slack, without prompt changes, saw about 14% fewer output tokens and fewer steps. None of that is a reason to point open-ended judgment work at this model. Anthropic says Opus 5.5 stays clearly stronger there, in their testing and in external tests.
What you have to change in the agent
Sonnet 5.5 is the first Sonnet with cyber safeguards of the kind Anthropic uses on its most capable models, because its cybersecurity capability is comparable to Opus 5. Higher-risk cyber tasks visibly fall back to Sonnet 5. Biology safeguards stay where Sonnet 5 left them. A refusal comes back as HTTP 200 with stop_reason: "refusal". The beta server-side fallback retries cyber and frontier-model declines on Sonnet 5. It does not retry biology, reasoning-extraction, or general-harms declines.
Five breaks will 400 an agent written for Sonnet 5. Thinking cannot be sent as disabled. The lowest setting is between_tools, and only at high effort or below. Forced tool choice errors. auto and none work. Thinking blocks from this model are not readable by Opus 5, Opus 5.5, Fable, or Mythos, and this model will not read theirs. It will read blocks from Sonnet 5, Opus 4.8, and Haiku 4.5. On API accounts created on or after August 31, 2026, editing the history before one of its thinking blocks returns a 400. On the Claude API and Google Cloud, computer_20251124 is rejected. Bedrock still accepts it. The advisor tool rejects Opus 4.8, Opus 4.7, and Sonnet 5 as advisors. Non-default temperature, top_p, or top_k also returns a 400. Effort levels are recalibrated, so a Sonnet 5 setting is not the same amount of thinking.
In practice
- Use claude-sonnet-5-5 at medium or high for well-scoped coding, bug fixes, spreadsheets, and slide drafts. That is where the cost charts sit in the cheap corner. Claude Code and the apps already default to medium. The Platform defaults to high.
- Do not move terminal work off Opus because Sonnet prints 70.6%. That score is max effort at $12.54. Opus's 66.4% is xhigh at $7.35. At medium, Sonnet is 28.8% and Opus is 57.6%.
- On FrontierCode, stop at xhigh. Sonnet is 52.1% at $1.59. Max is 46.2% at $20.78. Opus at medium is 54.6% at $0.80, above Sonnet's best.
- On Briefcase, medium and high are the Sonnet purchase: 1461 Elo at $1.64, and 1634 at $3.95. Max is 1811 at $29.19, against Opus at 1822 and $21.05.
- Do not pick Sonnet over GPT-6 Sol on the rate card. Both are $2 and $10, with a $0.20 cache read. Pick Sonnet when the deliverable is the knowledge-work scores above. Sol's GDPval is 1487 and its Briefcase max is 1483 at $2.67.
- A run that is mostly cache hits does not cost half of Opus. The $0.20 read is the same on both. The half-price is uncached input and output.
- Quote Terminal-Bench as two numbers from two sources. 70.6% at $12.54 is Anthropic's max-effort chart. 63.6% is Artificial Analysis, as listed by BenchLM.
- Before you point a Sonnet 5 agent at this id, switch thinking to between_tools, drop forced tool choice, leave temperature at the default, and drop Opus 4.8, Opus 4.7, and Sonnet 5 as advisors. A cyber refusal can fall back to Sonnet 5. A biology refusal will not.
Sources: Introducing Claude Sonnet 5.5, Claude Sonnet 5.5 on the Claude Platform, BenchLM, and Artificial Analysis. Effort prices are Anthropic's charts. The index of 56 and the $7.60 task are Artificial Analysis at max with fallback.