Model Releases 2026-09-28

Claude Sonnet 5.5: Same Per-Token Price, Up to 30% Cheaper Per Task -- and a Cyber Fallback Built Into the Model ID

Anthropic's new mid-tier model beats Opus 5.5 on Terminal-Bench 4.0 in Anthropic's own testing, trails it on harder coding work, and is the first Sonnet to ship with frontier-style cyber safeguards and anti-distillation classifiers.

Anthropic released Claude Sonnet 5.5 on September 28, 2026, the second model in the Claude 5.5 family after Opus 5.5. The pricing is unchanged from Sonnet 5 at $2 per million input tokens and $10 per million output ($0.20 for cache reads). Anthropic says the model generates output more than 30% faster and costs up to 30% less per task, and that the saving comes from using fewer tokens and fewer tool calls, not from a lower sticker price. It is available through Anthropic's API as claude-sonnet-5-5 and on AWS, Google Cloud and Microsoft Azure.

Anthropic's published comparisons, all from its own testing:

  • Terminal-Bench 4.0 (agentic coding): 70.6%, against 10.3% for Sonnet 5 and, as reported, 66.4% for Opus 5.5.
  • FrontierCode 1.1 (Main, max effort): 46.2%, against 54.4% for Opus 5.5. The larger model still leads on the hardest work.
  • CursorBench 4.0: 55.5%, against 34.1% for Sonnet 5.
  • OSWorld 2.1 (computer use): 80.1%. GDPval-AA v2.1 (knowledge work): 1844 Elo against 1449 for Sonnet 5.

Anthropic positions Sonnet 5.5 as the fast, cheaper complement to Opus 5.5, strongest on well-scoped everyday tasks: fixing bugs, and producing documents, slides and spreadsheets.

Two safety features change how the model behaves in practice. Sonnet 5.5 ships with cyber safeguards similar to Opus 5.5's, under which higher-risk cybersecurity tasks visibly fall back to Sonnet 5. And it is the first Sonnet with classifiers that block reasoning extraction, plus expanded "preserved thinking" so reasoning can't be separated from the account that generated it. That is a defence against competitors distilling its chain of thought.

The engineering implications: first, procurement should compare cost per completed task, not per-token rates, because the savings here are entirely behavioural and will vary with your workload. Second, security teams should expect one model ID to behave like two. Penetration-testing or exploit-analysis workflows may silently get Sonnet 5-level capability on the requests that matter most, so evaluate those paths separately. What remains uncertain is how much of the Terminal-Bench jump from 10.3% to 70.6% reflects general capability and how much reflects benchmark-specific fit. Independent evaluations over the next few weeks are the check.

Sonnet 5.5 cuts cost per task through efficiency rather than price, trails Opus 5.5 only on the hardest coding benchmark in Anthropic's own numbers, and introduces a model-internal fallback for high-risk cyber requests that security teams must test for explicitly.