Overview
Claude Sonnet 5.5, released September 28, 2026, is Anthropic's successor to Claude Sonnet 5. It targets teams that want a faster default for well-scoped coding, document work, and agent tasks. Anthropic reports stronger results than Sonnet 5, more than 30% faster output, and lower token use per task while keeping the same API token prices. The upgrade also changes several API behaviors, so production integrations need testing before a model ID swap.
The model is available as claude-sonnet-5-5 on the Claude API, Google Cloud, Microsoft Foundry, and Claude Platform on AWS, and as anthropic.claude-sonnet-5-5 on Amazon Bedrock. Its context window remains 1 million tokens, with up to 128,000 output tokens. Claude Opus 5.5 remains Anthropic's stronger option for complex, open-ended work that needs sustained judgment.
What's New
Better Results Across Coding and Knowledge Work
In Anthropic's published launch evaluations, Sonnet 5.5 scores 55.5% on CursorBench 4.0 against Sonnet 5's 34.1%, and 80.1% (partial) on OSWorld 2.1 against Sonnet 5's 57.0% (partial) for computer use. Its GDPval-AA v2.1 score rises from 1449 to 1844, nearly matching Opus 5.5's 1846 in that evaluation. These are benchmark results under Anthropic's reported settings; they do not establish the same improvement for every application.
The model is aimed at everyday bug fixing and well-defined tasks, but also produces more polished documents, slides, spreadsheets, and interfaces, according to Anthropic and its launch testers. Teams using Sonnet 5 for these tasks have a concrete reason to compare outputs on their own workload.
Faster Responses at the Same Token Rates
Anthropic says Sonnet 5.5 generates output more than 30% faster than Sonnet 5. Input and output prices remain $2 and $10 per million tokens, respectively; cache reads remain $0.20 per million. Anthropic estimates up to 30% less cost per task because Sonnet 5.5 typically uses fewer tokens for the same work. That is a workload estimate, not a reduction in the posted token rates. Actual savings depend on prompts, effort settings, caching, and output length.
New API Behavior and Safeguards
Adaptive thinking remains on by default as it was on Sonnet 5; Sonnet 5.5 changes the no-up-front-thinking path by replacing disabled with between_tools. It also supports per-message effort (beta), mid-conversation system messages, mid-conversation tool changes (beta), and a shorter 512-token minimum for prompt caching. Sonnet 5.5 adds cyber safeguards. On the Claude API, beta server-side fallback (fallbacks: "default") retries cyber and frontier_llm refusals on Sonnet 5; it does not retry bio, reasoning_extraction, or general_harms refusals. Teams handling security work should check how refusals and fallback affect their own workflows.
Compared With Sonnet 5
| Decision point | Sonnet 5.5 | Sonnet 5 |
|---|---|---|
| Published API input/output rate | $2 / $10 per million tokens | $2 / $10 per million tokens |
| Output generation | More than 30% faster in Anthropic's report | Baseline |
| Task cost | Up to 30% lower in Anthropic's testing | Baseline |
| Context window | 1 million tokens | 1 million tokens |
| CursorBench 4.0 | 55.5% | 34.1% |
Thinking without a thinking field |
Adaptive thinking on | Adaptive thinking on |
| Lowest thinking setting | between_tools |
disabled |
The strongest case for upgrading is better output quality and latency at the same posted rate. The strongest reason to delay is an integration that relies on settings Sonnet 5.5 rejects, or a workload where a changed safeguard or effort level has not been tested. For the hardest ambiguous tasks, compare Sonnet 5.5 with Opus 5.5 rather than assuming benchmark proximity makes them interchangeable.
Migration Guide
Change the Claude API model ID from claude-sonnet-5 to claude-sonnet-5-5, then check these release-specific changes before moving live traffic:
- Replace
thinking: {"type": "disabled"}withbetween_toolsif you need no up-front thinking.between_toolssupports low, medium, and high effort; higher effort levels require adaptive thinking. - Replace forced
tool_choicetypesanyandtoolwithauto. Where schema-valid tool arguments matter, use strict tools on supported platforms and validate the model's choice to call a tool. - Preserve the conversation history around Sonnet 5.5 thinking blocks. For accounts created on or after August 31, 2026 00:00 UTC, the Claude API, Amazon Bedrock, and Google Cloud enforce this conversation-prefix check by default; replaying an affected thinking block after earlier history changes returns a 400 error. Older accounts enforce it when
thinking.block_binding.prefix_mismatch_behavioris set. - On the Claude API or Google Cloud, replace the older
computer_20251124computer-use tool withcomputer_toolset_20260801. On Amazon Bedrock, Sonnet 5.5 continues to acceptcomputer_20251124. - If you use the advisor tool with Claude Opus 4.8, Claude Opus 4.7, or Claude Sonnet 5 as advisor, switch to a supported advisor; those pairings return a 400 error with a Sonnet 5.5 executor.
- Re-run latency, quality, and cost tests for your effort levels. Anthropic recalibrated them; the same setting does not imply the same thinking budget or behavior as Sonnet 5.
Applications that display text between tool calls should also inspect the new thinking progress-update blocks. Without an appropriate display setting, the interface can appear quiet while work continues. See Anthropic's full migration guide for supported parameters and platform details.
Who Should Upgrade / Who Should Wait
Upgrade after testing if Sonnet 5 powers routine coding agents, document production, or other high-volume, well-scoped work. The speed and task-efficiency claims are relevant where latency and cost per completed task matter more than token price alone. Start with a sample of real prompts and compare completion quality, elapsed time, and billed tokens.
Wait for integration fixes if your API code forces a specific tool, disables thinking with the old value, edits saved conversations, or uses the earlier computer-use tool on the Claude API or Google Cloud. Keep Opus 5.5 in the comparison for open-ended work requiring careful judgment; Anthropic says it remains stronger there despite close scores on some evaluations.




