A new Sonnet beside Opus
Anthropic launched Claude Sonnet 5.5 on 28 September as the second member of its 5.5 family. The Anthropic announcement positions it as the practical complement to Opus 5.5: a model for everyday coding, documents and repeatable professional tasks rather than the most judgement-intensive work. Haiku 5.5 is promised for a later date, not included in this release.
Anthropic says Sonnet 5.5 generates output more than 30 per cent faster than Sonnet 5. Its listed API rates stay at US$2 per million input tokens and US$10 per million output tokens, with cache reads at US$0.20. The company nevertheless claims a task can cost up to 30 per cent less because the model often uses fewer tokens to reach a result.
Those are vendor measurements, not a universal price reduction. A task that needs repeated tool calls, long context or a human correction can have very different economics. Buyers should measure accepted work, elapsed time and total billed tokens on their own prompts before replacing an existing model.
What the published tests suggest
Anthropic reports a 70.6 per cent score for Sonnet 5.5 on Terminal-Bench 4.0, versus 10.3 per cent for Sonnet 5 in its comparison. It also reports stronger performance on CursorBench, a coding-agent evaluation based on ambiguous multi-file tasks. Such numbers indicate a substantial change in the company’s test setup, but they do not forecast success on every private repository.
The company also points to gains in occupational knowledge work, chart interpretation and computer use. Its examples from early users describe fewer tool calls and quicker iteration. These observations are useful signals, though they come from selected testers and tasks. Teams should keep their own held-out set, including the awkward cases that caused problems with Sonnet 5.
Effort level matters to the comparison. Claude apps and Claude Code default to Medium, while the Claude Platform defaults to High. Higher effort may improve a difficult answer but increase token use and waiting time. A fair evaluation runs each model with the settings an organisation will actually deploy, not just its best published score.
Availability and migration details
Sonnet 5.5 is available through the Claude Platform under the identifier claude-sonnet-5-5, and Anthropic says it is also available via Amazon Web Services, Google Cloud and Microsoft Azure. That breadth may simplify procurement for teams already using one of those environments. Regional availability, quotas and contract terms should still be confirmed with the relevant provider.
Developers who previously disabled thinking on Sonnet need to review a behavioural change. Anthropic says the new between_tools setting is the closest way to avoid up-front extended thinking. It can still return short progress updates between tool calls as thinking blocks. Applications that parse responses or transfer conversation state between accounts should test the migration guide before switching production traffic.
The model supports zero data retention according to Anthropic, but that statement should be read alongside the terms of the chosen access route. Logging, tool outputs and application storage can sit outside a model provider’s retention setting. Australian teams working with customer records should map where each component sends and stores data.
Safety is part of this release
Anthropic says the model’s cybersecurity capabilities are closer to its larger models, so Sonnet 5.5 ships with safeguards and fallbacks similar to those used for Opus 5.5. Biology safeguards remain aligned with Sonnet 5. The company describes these as controls for narrow, high-risk requests, rather than a restriction on ordinary software development or life-sciences work.
Its behavioural audit covered roughly 1,850 scenarios. Anthropic says Sonnet 5.5 improved or matched the previous Sonnet on most measures, while Opus 5.5 remained slightly stronger overall. A lab-run audit cannot find every failure. An organisation deploying agentic features still needs permissions, review gates and monitoring around tools that can change files, send information or transact.
For coding agents, the operational test is not whether a model can write an impressive diff. It is whether the change stays in scope, passes tests and can be reviewed. Anthropic notes that higher effort sometimes produced extra edits or timeouts on one benchmark. More reasoning is not automatically a safer default for a bounded engineering task.
Where it is worth testing
The strongest adoption case is a queue of common tasks for which the current Sonnet model is accurate enough but slow or expensive in aggregate. Bug fixes, document production and structured analysis fit the positioning. Start with representative tasks and record quality, latency, retries, cache use and the time a person spends checking the output.
Keep harder tasks on Opus where the extra capability justifies its cost. A routing policy can send a first attempt to Sonnet 5.5, then escalate a failure or an ambiguous case. Such a policy needs clear failure signals; otherwise it simply adds another model call and obscures responsibility for the final answer.
Sonnet 5.5 gives developers a genuine new option, with a clear release date, identifier and price. Its value will be the reduction in cost and waiting time per accepted result, not a benchmark headline taken in isolation. Teams can establish that with a controlled pilot before changing default models for everyone.