Last updated: October 5, 2026

If you are choosing between Claude Opus 5.5 and Claude Sonnet 5.5 for coding or agent work, the sticker price is only half the story. Sonnet 5.5 costs half as much per token and even beats the flagship on one agentic coding benchmark, yet Opus 5.5 still leads most long-horizon tests. This guide gives you a simple routing rule so you stop paying flagship prices for routine work.

Quick answer: Sonnet 5.5 ($2/$10) beats Opus 5.5 on Terminal-Bench at half the price. Use Sonnet for scoped coding; save Opus 5.5 ($4/$20) for long agent runs.

What changed with Opus 5.5 and Sonnet 5.5?

Anthropic released Opus 5.5 on 22 September 2026 at $4 per million input tokens and $20 per million output, a 20% cut from the previous Opus price, and followed with Sonnet 5.5 on 28 September 2026 at $2/$10 — the same price as Sonnet 5 (Temperature, 29 September 2026). Both models offer a 1 million-token context window and 128K maximum output. In plain terms: Opus is positioned for long-running agentic work, Sonnet for the best mix of speed and intelligence at half the price.

Claude Opus 5.5 vs Sonnet 5.5 at a glance

Use this table as the shortlist before you read the benchmarks. Prices are Anthropic list prices as reported at launch; always re-check the vendor page before budgeting.

Item Sonnet 5.5 Opus 5.5
Release 28 Sep 2026 22 Sep 2026
API price (input/output per 1M) $2 / $10 $4 / $20
Cache read (per 1M) $0.20 $0.20
Context / max output 1M / 128K 1M / 128K
Best for Scoped coding, terminal work, writing, documents Long, ambiguous, multi-file agent runs

Sources: launch pricing and dates from Temperature (29 Sep 2026); cache and context figures from the model comparison by Aleksei Aleinikov (Sep 2026). Desk-validated, not hands-on tested by us.

Where does Sonnet 5.5 actually beat Opus 5.5?

On Terminal-Bench 4.0, an agentic command-line coding test, Anthropic’s published table gives Sonnet 5.5 70.6% against Opus 5.5’s 66.4% (Anthropic launch table, reported 29 Sep 2026). That is the one headline row the cheaper model wins outright — and it matters because terminal work is what coding agents do all day. Caveat: these are vendor-published results and had not been independently reproduced at the time of reporting, so treat the gap as directional, not decisive.

Where does Opus 5.5 still lead?

Everywhere the work gets longer and less well defined. In the same Anthropic table, Opus 5.5 leads FrontierCode 1.1 at max effort (54.4% to 46.2%), CursorBench 4.0 (57.8% to 55.5%) and OSWorld 2.1 computer use (81.8% to 80.1%), while GDPval-AA v2.1 knowledge work is a near tie (1,846 to 1,844 Elo). On the separate SWE-bench Pro leaderboard, Opus 5.5 leads at 89.9% to Sonnet 5.5’s 81.3% as of 2 October 2026 (BenchLM, October 2026) — useful for long repository work, but splits and scaffolds differ between published rows, so do not buy on that number alone. Our GPT-6.1 Sol guide shows the same pattern at OpenAI: the second-tier model wins the value argument, the flagship wins the hardest runs.

What does each model really cost per task?

List price says 2x; a real agent loop says less. Cache reads cost the same $0.20 per million on both models, and agent loops re-read cached context every turn. In a worked per-turn example (60K cache reads, 3K cache writes, 1.5K output), Sonnet 5.5 totals about $0.0345 against Opus 5.5’s $0.0570 — a 1.65x premium, not 2x (worked example in Aleinikov, Sep 2026). Independent agent data tells the same effort story: on the Artificial Analysis Coding Agent Index v1.5 as of 1 October 2026, Sonnet 5.5 at max effort scores 68.4 but costs $14.19 per task and takes 87.4 minutes, while at xhigh it scores 62.9 for $3.33 (Preuve AI summary of Artificial Analysis, Oct 2026). The savings live at low and medium effort; pushing Sonnet to max converges on Opus pricing.

How to choose in 5 steps (our routing heuristic)

This is the original element of this guide: a step-count rule we derived from the published scores, effort curves and per-turn maths above. It is desk-validated, not hands-on tested — run your own evals before locking it in.

  1. Count the steps. Under ~10 agent steps, one file, a clear spec? Start on Sonnet 5.5 at medium effort.
  2. Check ambiguity. Vague requirements, architecture decisions, or a wrong answer that costs more than the tokens? Start on Opus 5.5.
  3. Watch the file count. Changes that must stay consistent across 5+ files lean Opus; single-file fixes and CI/terminal jobs lean Sonnet.
  4. Set effort before you switch models. Raise Sonnet from medium to high only when your evals prove it; avoid Sonnet at max (cost converges on Opus without Opus judgement).
  5. Escalate by summary, not by switching mid-chat. If Sonnet stalls, summarise the task state and open a fresh Opus conversation (see the routing trap below).

For migration context on Google’s side, see our Gemini Skills migration guide.

What is the routing trap nobody mentions?

The two models cannot read each other’s thinking blocks. According to the Claude platform documentation summarised by Aleinikov (Sep 2026), Sonnet 5.5 does not read Opus 5.5 thinking, and no other model reads Sonnet 5.5 thinking; on a model switch the block is silently dropped while the request still succeeds. So a per-turn router quietly throws away the reasoning it just paid for. Route per conversation instead, or keep Sonnet in charge and consult Opus through the advisor tool. If you build routing, log the dropped-block signals so the loss is visible. This mirrors the safety-first setup in our OpenClaw AI agent review.

How do you set this up in Claude Code and the API?

  1. In Claude Code, pick the model with /model and the effort with /effort; note both models default to medium effort in Claude Code unless you change it (Artificial Analysis data via Preuve, 1 Oct 2026).
  2. On the API, set effort explicitly in your request — Sonnet 5.5 defaults to high and Opus 5.5 to medium, so an unset A/B test is not a fair comparison.
  3. Keep histories append-only and avoid editing earlier messages, system prompts or tool lists mid-session; newer accounts enforce this strictly.
  4. Turn on prompt caching for long agent loops — cache reads are where the two models cost the same, and caching is what shrinks the real premium.
  5. Re-baseline after any model swap: effort levels were recalibrated for 5.5, so last month’s setting is not this month’s setting.

Should you wait for Haiku 5.5 or switch away?

Anthropic says a smaller Haiku 5.5 follows “in the coming weeks” with no price announced yet (Temperature, 29 Sep 2026) — label that unconfirmed until it ships. For high-volume, latency-sensitive work it may undercut both models here. If pure cost per coding task is your metric, note that the same Artificial Analysis index puts Codex with GPT-6.1 Sol at xhigh at 62.9 for $1.04 per task (Preuve, Oct 2026) — the current value benchmark to beat. Do not wait idle, though: if Sonnet 5.5 at medium clears your quality bar today, ship it and re-run the routing check when Haiku lands.

Frequently asked questions

Which Claude model is cheaper, Opus 5.5 or Sonnet 5.5?

Sonnet 5.5, at $2/$10 per million tokens against Opus 5.5’s $4/$20 (Anthropic launch pricing, Sep 2026). In cache-heavy agent loops the real premium narrows to roughly 1.65x per turn because cache reads cost $0.20 on both.

Which Claude model is better for coding?

For scoped, well-defined coding and terminal work, Sonnet 5.5 — it leads Terminal-Bench 4.0 at 70.6% in Anthropic’s published table. For long, multi-file, ambiguous runs, Opus 5.5 leads FrontierCode, CursorBench and SWE-bench Pro.

Should I run Sonnet 5.5 at max effort?

Usually no. Artificial Analysis data (1 Oct 2026) shows max effort multiplying cost and time for a modest score gain; Anthropic’s own framing puts Sonnet’s savings at lower effort settings. Prove it on your evals first.

Can I switch between Sonnet 5.5 and Opus 5.5 mid-conversation?

You can, but the next model will not see the previous model’s thinking blocks — they are silently dropped per the platform docs. Route per conversation, or escalate with a written summary into a fresh Opus chat.

Sources and methodology

Claims rest on Anthropic’s published launch figures as reported by Temperature (29 Sep 2026), the model and routing analysis by Aleksei Aleinikov (Sep 2026), the Artificial Analysis Coding Agent Index v1.5 summary by Preuve AI (1 Oct 2026), and the BenchLM SWE-bench Pro leaderboard (2 Oct 2026). Method: desk comparison of vendor-published benchmarks against one independent index, verified 5 October 2026; vendor scores are labelled as such and were not hands-on reproduced by us. Benchmark gaps are correlations across different harnesses and splits, not controlled head-to-head findings.

Have a burning question about this topic?
Feel free to email us at contact@openaimaster.ai — we are happy to help!