Claude Sonnet 5 Is Cheaper Per Token and Sometimes Pricier Per Task Than Opus

Claude Sonnet 5 Is Cheaper Per Token and Sometimes Pricier Per Task Than Opus

6 min readJuly 1, 2026

Quick verdict

Anthropic launched Claude Sonnet 5 as the new default mid-tier model, and on paper it is a clear upgrade: a 1M-token context window, strong coding and tool-use scores, and promotional pricing of $2 per million input tokens and $10 per million output through the end of August. The catch is subtle. Sonnet 5 talks and works more than Sonnet 4.6 did, so it spends more tokens and takes more agentic turns to finish a job. Artificial Analysis measured a solved task at about $2.29, roughly double Sonnet 4.6 and about 15% more than Opus 4.8, even though Sonnet's per-token price is lower. Cheaper per token, sometimes pricier per task. That gap is the whole story.

What actually shipped

Anthropic's own account called Sonnet 5 its "most agentic Sonnet yet," with planning, browser and terminal tool use, and autonomous execution that until recently needed larger models. The developer account added the specifics builders care about: a 1M-token context window, frontier-level coding at Sonnet pricing, and immediate availability as the default model in Claude Code for Pro users, across the API, and in Managed Agents. Anthropic kept the standard list price at $3 input and $15 output per million tokens, with the $2/$10 promo running through the end of August.

The model also picked up a fifth effort setting, xhigh, matching Opus 4.8's ladder of max, xhigh, high, medium, and low. Cache pricing follows the usual pattern: a 25% premium on cache writes and a 90% discount on cache hits, with a five-minute window. Alongside the model, Anthropic shipped Claude Desktop for Linux in beta on Ubuntu and Debian, plus a batch of Managed Agents updates like streaming session deltas, webhook events, and a token and tool observability tab.

The benchmark numbers back the "real upgrade" framing:

BenchmarkSonnet 5Reference
SWE-bench Pro63.2%up from Sonnet 4.6
Terminal-Bench 2.180.4%+9 over 4.6
Humanity's Last Exam (with tools)57.4%+10 over 4.6
OSWorld-Verified81.2%agentic desktop tasks
CursorBench (Cursor)57%vs 49% for Sonnet 4.6
AA Intelligence Index53+6 over 4.6, #5 overall

On the Artificial Analysis Intelligence Index, Sonnet 5 lands at 53, up six points, which puts it fifth overall and roughly level with GPT-5.5 on high reasoning, still behind Opus 4.7 and 4.8. Cognition reported that on its FrontierCode Extended eval, Sonnet 5 scored 53.8% with a 57.6% pass rate, actually ahead of Opus 4.8 in their harness. Cline found Opus 4.8-level Terminal-Bench performance for less than half the cost, plus better resistance to prompt-injection hijacks, which matters for anyone letting an agent run terminal or browser commands unattended.

Why the per-task cost is the real number

Here is where the list price and the invoice part ways. Artificial Analysis found Sonnet 5 used about 69,000 output tokens per task on average, roughly 40% more than Sonnet 4.6. It also took around three times as many agentic turns on their agentic benchmarks, and max effort ran up about six times the turns of low effort. More tokens and more turns mean that a lower price per token does not translate into a lower price per finished task. Their measurement put a solved Intelligence Index task at about $2.29, close to double Sonnet 4.6 and about 15% above Opus 4.8.

Simon Willison flagged a second wrinkle: the new tokenizer makes Sonnet 5 about 1.4 times more expensive for English text and 1.33 times for Spanish, while Simplified Mandarin stays roughly flat. So even the per-token comparison understates the real cost for English-heavy work, because the same paragraph now counts as more tokens. This is why the loudest critics, including Theo and several others, kept pointing out that Sonnet 5 can cost more than Opus 4.8 or even Fable on the tasks they actually ran, despite the friendlier sticker price.

The counterpoint from builders is fair too. For long-running agents, parallel workflows, and production coding loops, a stronger cheap-tier default is exactly what teams want, and static per-token benchmarks undersell reliability gains on long-horizon work. The honest read is that Sonnet 5 is a production-friendly release rather than a headline frontier jump. Whether it saves you money depends entirely on whether you measure per token or per task.

Why it matters

The old rule of thumb was simple: reach for Sonnet to save money, reach for Opus when the job is hard. Sonnet 5 blurs that. If you are running agentic coding where the model chews through tokens and turns, the "cheap" default can quietly cost more than the flagship. The fix is the same discipline we described in our look at how enterprises cut AI spend by routing instead of rationing: pick the model per task, watch cost per completed job rather than cost per token, and cap effort levels when max is not buying you a better answer.

The ecosystem moved fast on this. Cursor, Perplexity, Devin, Cline, Factory's Droid, VS Code, and Agent Arena all added Sonnet 5 within a day, several with their own temporary discounts, which tells you the market sees it as the new workhorse default even where enthusiasm is mixed. If you are choosing a coding agent around it, our Claude Code vs Cursor vs Codex comparison and our best AI models for coding roundup are the places to start.

Video: Claude Sonnet 5 first look

A walkthrough of what Sonnet 5 changes for coding and agent workflows.

FAQ

Is Claude Sonnet 5 actually cheaper than Opus 4.8?

Per token, yes, especially during the $2/$10 promo. Per finished task, not always. Artificial Analysis measured a solved benchmark task at about $2.29 for Sonnet 5, roughly 15% more than Opus 4.8, because Sonnet 5 uses more output tokens and more agentic turns. Measure cost per completed job, not per token.

What is the promotional pricing and how long does it last?

Standard pricing is $3 per million input tokens and $15 per million output. The promo rate is $2 input and $10 output, running through the end of August. After that it reverts to list price unless Anthropic extends it.

Does Sonnet 5 beat Opus 4.8 on benchmarks?

On most broad intelligence measures, no. It sits around fifth on the Artificial Analysis Intelligence Index, behind Opus 4.7 and 4.8. On some coding evals it does pull ahead: Cognition's FrontierCode Extended put Sonnet 5 above Opus 4.8 in their harness. It is a coding and agent specialist more than an across-the-board frontier model.

Should I switch my coding agent to Sonnet 5?

If you run agentic coding at volume, test it against your own workload and track the total cost of a completed task, not the per-token rate. It is already the default in Claude Code for Pro users and supported across Cursor, Cline, Devin, and others. See our coding agent comparison for the tradeoffs.

Sources

Further reading

Try all the models mentioned in this article

Admix gives you GPT-5, Claude, Gemini, and 350+ AI models in one app. Compare responses side by side. Free to start.

Start free on Admix

Related articles