Anthropic has released Claude Sonnet 5.5, the second model in its Claude 5.5 family. On several headline benchmarks it comes within a point or two of the company's flagship Opus 5.5, at half the per-token price. The September 28 launch keeps Sonnet 5's rates of $2 per million input tokens and $10 per million output tokens. Anthropic says the model generates output more than 30% faster and can cut the cost of a typical task by up to 30%.

The release comes six days after Opus 5.5 debuted and OpenAI answered with its GPT-6 Sol and Luna models. Price competition at the top of the market has become intense.

Benchmarks: Closing the Gap With Opus

Anthropic's reported results show a big jump over Sonnet 5 and near-parity with Opus 5.5 on agentic and knowledge-work tests:

  • Terminal-Bench 4.0: 70.6%, up from 10.3% for Sonnet 5.
  • GDPval-AA v2.1, which measures economically valuable knowledge work: an Elo of 1844, against 1449 for Sonnet 5 and 1846 for Opus 5.5.
  • OSWorld 2.1, for computer use: 80.1%, compared with 81.8% for Opus 5.5.
  • Humanity's Last Exam with tools: 64.5%, against 67.7% for Opus 5.5.

On FrontierCode 1.1, reported figures put Sonnet 5.5 at 52.1% at its highest effort setting, above GPT-6 Sol's 49.3% and just behind Opus 5.5. Anthropic also says Sonnet 5.5 is the first Sonnet model to beat Pokémon Red using only screenshots. That is an informal but widely followed test of long-horizon planning.

Independent benchmarker Artificial Analysis put Sonnet 5.5 at 56 on its Intelligence Index, two points behind Opus 5.5 at maximum effort.

Effort Levels and the Real Cost Story

The per-token price has not changed, so any savings have to come from the model using fewer tokens. Sonnet 5.5 offers a range of effort levels: Low and Medium for routine work, and High, Xhigh and Max for harder problems that need longer reasoning and more checking. Anthropic says that on several benchmarks, Sonnet 5.5 at Low or Medium effort beats Sonnet 5's best score at about a tenth of the cost per task.

The picture changes at the top of the dial. Artificial Analysis found that at maximum effort, Sonnet 5.5 generated about 193,000 output tokens per question on average. That was the highest of any configuration it tested. It worked out to about $7.60 per Intelligence Index task, roughly 50% more than Sonnet 5 and more than Opus 5.5 at about $5.98.

  • Routine and agentic work: Low and Medium effort is likely to deliver the advertised savings.
  • Hard reasoning: Max effort can become more expensive than the flagship model.
  • Real-world signals are positive. In one published code-review comparison, Sonnet 5.5 cost about $0.46 per run versus $1.16 for Sonnet 5, and it produced far fewer trivial comments.

Safety and Anti-Distillation Measures

Anthropic reports better scores than Sonnet 5 on its automated behavioural alignment audit. The model ships with cybersecurity safeguards similar to Opus 5.5's, which Anthropic says leave routine debugging unaffected. Its biology safeguards match Sonnet 5's. The company also added new distillation protections designed to stop attackers from extracting the model's reasoning to train copycat systems, a growing concern as open-weight rivals close the gap.

Sonnet 5.5 is available through the Claude Platform, Amazon Web Services, Google Cloud and Microsoft Foundry. Early customer reports from companies including Zendesk, Atlassian and Epic Games describe faster agent runs and fewer failed tool calls.

Why It Matters

Sonnet 5.5 continues a trend that is reshaping AI research economics: frontier-class capability is moving down to the mid tier within weeks, not years. When a mid-priced model scores within two Elo points of the flagship on GDPval, most production workloads no longer need the largest model. The main exceptions are open-ended problems that call for sustained judgment, where Anthropic still recommends Opus 5.5.

The Artificial Analysis findings also mark a shift in how models should be evaluated. Per-token price is becoming a weak guide to real cost. Token efficiency, effort settings and task success rates now matter more than the rate card. Developers should benchmark on their own workloads at several effort levels before switching. The cheapest model per token is not always the cheapest model per answer.

Anthropic has indicated that Haiku 5.5 is next in the family. That will test whether the same efficiency gains carry over to the smallest, fastest tier.

Sources