VJOURNAL

AI • Global Desk • October 01, 2026

Claude Sonnet 5.5 challenges Opus on cost, but coding effort changes the result

Sonnet 5.5 brings lower token tariffs into Anthropic’s new lineup. Shared context limits and an unusual coding result make the choice more subtle than price alone.

AI-assisted conceptual still life of small and large anonymous mechanical assemblies with an unmarked balance arm on walnut; no real Anthropic event.

Answer in brief

Sonnet 5.5 brings lower token tariffs into Anthropic’s new lineup. Shared context limits and an unusual coding result make the choice more subtle than price alone.

Evidence cutoff: 5 sources
Sonnet 5.5 launched on 28 September after Opus 5.5’s 22 September announcement.
Both document one million context tokens and 128,000 output tokens, with different ordinary API tariffs.
Sonnet’s FrontierCode score is lower at maximum effort than at xhigh under the published evaluation rules.

A cheaper sibling changes the default question

Anthropic introduced Claude Sonnet 5.5 on 28 September, following Opus 5.5’s official announcement on 22 September. The new pairing makes Claude model costs a practical question for coding and document teams: which tasks warrant the larger model, and which can be completed reliably with the less expensive sibling?

Anthropic positions Sonnet for defined everyday work and Opus for complex work requiring sustained judgment. That is a provider description, rather than a result from VJOURNAL testing. The useful implication is to examine the shape of a task before selecting a model: a precise bug fix and an ambiguous architecture review demand different kinds of work.

Context is shared; the ordinary tariffs are not

The model documentation checked on 1 October lists the same one million context tokens and 128,000 output tokens for both. Sonnet’s ordinary input and output tariffs are half Opus’s. Cache reads are equal, however, and five-minute cache writes have their own rates. The table keeps those categories separate rather than compressing them into a blended price.

Equal capacity does not demonstrate equal recall across a long repository. Nor does a cheaper input token establish a lower invoice for the whole job. Repeatedly supplied context, generated reasoning and tool iterations can dominate a session. Teams should record their actual cache mix and accepted results before applying these tariffs to a forecast.

Anthropic model documentation, checked 1 Oct 2026. Standard API rates in USD per 1M tokens; cache writes below use the five-minute rate. Context and output are separate limits.
MeasureClaude Sonnet 5.5Claude Opus 5.5
Input USD / 1M tokens$2$4
Output USD / 1M tokens$10$20
Cache read USD / 1M$0.20$0.20
Cache write USD / 1M$2.50$5
Context tokens1,000,0001,000,000
Output tokens128,000128,000
API default effortHigh (high)Medium (medium)
Input modalitiesText and imagesText and images

The surprising coding result depends on effort

Anthropic’s 28 September announcement reports Sonnet at 52.1% on FrontierCode v1.1 Main with xhigh effort, falling to 46.2% at max. Its footnote explains that more review activity sometimes caused a timeout or edits outside the requested scope. Opus’s reported max figure is 54.4%. These are provider-reported configurations, not a newly run editorial contest.

The benchmark asks whether a change can be merged without human edits. Doing more work can therefore lose points if it broadens the task unnecessarily. The result is a concrete reason to specify the acceptance boundary and choose an effort setting deliberately, instead of treating the largest reasoning budget as an automatic quality improvement.

Anthropic Sonnet announcement, 28 Sep 2026; checked 1 Oct. FrontierCode v1.1 Main, provider-reported percentages. Exact run dates Unknown; configuration changes matter.
ModelReasoning settingFrontierCode v1.1 Main (%)
Claude Sonnet 5.5Very high (xhigh)52.1%
Claude Sonnet 5.5Maximum (max)46.2%
Claude Opus 5.5Maximum (max)54.4%

Defaults and published scores need their own labels

Sonnet’s API documentation lists high as the default effort; its announcement says the Claude apps and Claude Code default to medium. Opus’s API documentation lists medium. Comparing two sessions without recording where they ran can therefore compare different reasoning budgets even when the user believes they selected ordinary settings.

Artificial Analysis also publishes an independent Terminal-Bench 4.0 evaluation using its own common harness. Those measurements are useful external evidence, but they should retain their settings and source identity. A vendor launch figure and an independent terminal result are different observations; swapping them between tables would obscure rather than strengthen the comparison.

Use a routing rule that can be checked

Our editorial inference is to begin with a defined task, a written stopping condition and a named effort level. Evaluate routine fixes and document revisions separately from open architectural decisions. If Sonnet meets the acceptance bar, its tariff advantage can matter. If it needs repeated rescue, the higher-capability route deserves a direct comparison.

Track complete cost, elapsed time and reviewer corrections, including failed attempts. Reassess after a model update rather than assuming the same split forever. At the 1 October cutoff, the release dates, capacity and ordinary prices are documented; the cheapest configuration for a particular team remains dependent on its workload and review standards.

Questions and answers

Is Sonnet always half the cost of Opus?

Its ordinary input and output token tariffs are half Opus’s, but cache reads cost the same. Different token use, retries and review time can change the final task bill.

Do their context limits differ?

The model documentation checked on 1 October lists one million context tokens and 128,000 output tokens for both. These are separate capacity specifications.

Why can maximum effort score lower?

Anthropic says extra review activity sometimes caused timeouts or changes beyond the task’s scope in FrontierCode. That test rewards changes accepted without human editing.