Answer in brief
Sonnet 5.5 brings lower token tariffs into Anthropic’s new lineup. Shared context limits and an unusual coding result make the choice more subtle than price alone.
A cheaper sibling changes the default question
Anthropic introduced Claude Sonnet 5.5 on 28 September, following Opus 5.5’s official announcement on 22 September. The new pairing makes Claude model costs a practical question for coding and document teams: which tasks warrant the larger model, and which can be completed reliably with the less expensive sibling?
Anthropic positions Sonnet for defined everyday work and Opus for complex work requiring sustained judgment. That is a provider description, rather than a result from VJOURNAL testing. The useful implication is to examine the shape of a task before selecting a model: a precise bug fix and an ambiguous architecture review demand different kinds of work.
Context is shared; the ordinary tariffs are not
The model documentation checked on 1 October lists the same one million context tokens and 128,000 output tokens for both. Sonnet’s ordinary input and output tariffs are half Opus’s. Cache reads are equal, however, and five-minute cache writes have their own rates. The table keeps those categories separate rather than compressing them into a blended price.
Equal capacity does not demonstrate equal recall across a long repository. Nor does a cheaper input token establish a lower invoice for the whole job. Repeatedly supplied context, generated reasoning and tool iterations can dominate a session. Teams should record their actual cache mix and accepted results before applying these tariffs to a forecast.
| Measure | Claude Sonnet 5.5 | Claude Opus 5.5 |
|---|---|---|
| Input USD / 1M tokens | $2 | $4 |
| Output USD / 1M tokens | $10 | $20 |
| Cache read USD / 1M | $0.20 | $0.20 |
| Cache write USD / 1M | $2.50 | $5 |
| Context tokens | 1,000,000 | 1,000,000 |
| Output tokens | 128,000 | 128,000 |
| API default effort | High (high) | Medium (medium) |
| Input modalities | Text and images | Text and images |
The surprising coding result depends on effort
Anthropic’s 28 September announcement reports Sonnet at 52.1% on FrontierCode v1.1 Main with xhigh effort, falling to 46.2% at max. Its footnote explains that more review activity sometimes caused a timeout or edits outside the requested scope. Opus’s reported max figure is 54.4%. These are provider-reported configurations, not a newly run editorial contest.
The benchmark asks whether a change can be merged without human edits. Doing more work can therefore lose points if it broadens the task unnecessarily. The result is a concrete reason to specify the acceptance boundary and choose an effort setting deliberately, instead of treating the largest reasoning budget as an automatic quality improvement.
| Model | Reasoning setting | FrontierCode v1.1 Main (%) |
|---|---|---|
| Claude Sonnet 5.5 | Very high (xhigh) | 52.1% |
| Claude Sonnet 5.5 | Maximum (max) | 46.2% |
| Claude Opus 5.5 | Maximum (max) | 54.4% |
Defaults and published scores need their own labels
Sonnet’s API documentation lists high as the default effort; its announcement says the Claude apps and Claude Code default to medium. Opus’s API documentation lists medium. Comparing two sessions without recording where they ran can therefore compare different reasoning budgets even when the user believes they selected ordinary settings.
Artificial Analysis also publishes an independent Terminal-Bench 4.0 evaluation using its own common harness. Those measurements are useful external evidence, but they should retain their settings and source identity. A vendor launch figure and an independent terminal result are different observations; swapping them between tables would obscure rather than strengthen the comparison.
Use a routing rule that can be checked
Our editorial inference is to begin with a defined task, a written stopping condition and a named effort level. Evaluate routine fixes and document revisions separately from open architectural decisions. If Sonnet meets the acceptance bar, its tariff advantage can matter. If it needs repeated rescue, the higher-capability route deserves a direct comparison.
Track complete cost, elapsed time and reviewer corrections, including failed attempts. Reassess after a model update rather than assuming the same split forever. At the 1 October cutoff, the release dates, capacity and ordinary prices are documented; the cheapest configuration for a particular team remains dependent on its workload and review standards.
Questions and answers
Is Sonnet always half the cost of Opus?
Its ordinary input and output token tariffs are half Opus’s, but cache reads cost the same. Different token use, retries and review time can change the final task bill.
Do their context limits differ?
The model documentation checked on 1 October lists one million context tokens and 128,000 output tokens for both. These are separate capacity specifications.
Why can maximum effort score lower?
Anthropic says extra review activity sometimes caused timeouts or changes beyond the task’s scope in FrontierCode. That test rewards changes accepted without human editing.
