Claude Haiku 5.5 Is Cheaper Per Token and More Expensive Per Task

Claude Haiku 5.5 Is Cheaper Per Token and More Expensive Per Task

6 min readOctober 9, 2026

Quick verdict

Anthropic dropped Claude Haiku 5.5 as its cheapest and fastest small model, priced at $0.10 per million input tokens and $0.50 per million output, and on the same day it halved Sonnet 5.5 cache reads. The per-token numbers look great, roughly 75% below Haiku 4.5 on average. Then the independent testers ran it and found something awkward: Haiku 5.5 reasons through so many more steps than its predecessor that on every shared benchmark it costs more per task than Haiku 4.5 did. That is the whole lesson in one model. The price on the pricing page is not what you pay. What you pay is price per token times tokens per task, and a cheaper, chattier model can quietly lose you money.

What shipped

Haiku 5.5 is Anthropic's small, high-volume model, meant for summarization, classification, live support, browser use, and as a cheaper sub-agent next to Opus and Sonnet 5.5. The headline specs are real improvements:

  • Priced at $0.10 / $0.50 per million input / output tokens, matching GPT-6 Luna and about 75% under Haiku 4.5 on average
  • A 1M-token context window and up to 128K output tokens
  • An adjustable effort setting, so you can turn the reasoning down when you do not need it
  • A big jump on coding evals: 90.4% on Vibe Code Bench, third overall, and 1587 on Code Arena's WebDev board, up 257 points over Haiku 4.5

Alongside it, Anthropic halved Sonnet 5.5 cache reads to $0.10 per million tokens, with input at $2 and output at $10. The company estimates that makes most agentic work about 20% cheaper, and Claude Code usage limits are unchanged. Taken together it reads as a straightforward price cut across the small and mid tiers.

The catch: more steps, bigger bill

Vals AI measured what the model actually does, not just what it costs per token, and the gap is the story. On their Legal Research task, Haiku 5.5 took 59 steps where Haiku 4.5 took 17. More steps means more tokens, and more tokens at any price is more money. The result is that Haiku 5.5 costs more per test than Haiku 4.5 on every benchmark the two share, even though each token is far cheaper.

Two more wrinkles make the real cost harder to predict. The first is a context cliff: Vals reports the price rises 5x once you go past 100k tokens of context, so the cheap rate is really the short-context rate and long documents do not get it. The second is scope. Haiku 5.5 scores 54.3% on the Vals Index, which lands it at #16. It is strong on coding and, per Roboflow, cheaper than GPT-6 Luna on high-effort vision work, but it is a small model, not a flagship, and one widely shared test clocked it at roughly 12x the cost of Luna on the same job.

None of this makes Haiku 5.5 a bad model. It is genuinely fast, cheap on short prompts, and good at code. The point is narrower: you cannot read its real cost off the pricing page, because how hard the model decides to think is now a variable you are paying for.

How to actually keep the cost down

This is the same trap we flagged with other "cheaper" launches. The useful number is cost per finished task, not cost per million tokens, and the two can point in opposite directions. We dug into exactly this gap in our write-up on the cost-per-task benchmark, where the cheapest-looking model is routinely not the cheapest to actually use.

A few practical moves follow from that:

  • Use Haiku 5.5's effort setting. If a task does not need deep reasoning, turning effort down is the difference between its token price being a bargain and a bill.
  • Watch the 100k-token line. Below it Haiku is cheap; above it the 5x jump can make a mid-tier model the better buy.
  • Measure on your own workload. The model that wins on Vibe Code Bench can still lose on your legal-research pipeline, because step counts depend on the task.

The broader habit is the one worth building: route each job to the model that clears the bar for the least real money, and keep several within reach so you can switch when the math changes. That is why we keep writing about how to stop overpaying for AI, and why running every model through one AI aggregator beats committing to any single lab's price list.

Video: what changed in Haiku 5.5

A quick walkthrough of the new pricing, the context window, and the effort setting, if you want the launch in short form.

FAQ

How much does Claude Haiku 5.5 cost?

$0.10 per million input tokens and $0.50 per million output, which matches GPT-6 Luna and is about 75% below Haiku 4.5 on average. Note that Vals reports the rate rises 5x beyond 100k tokens of context, so long-context work is not priced at the headline number.

Is Haiku 5.5 cheaper than Haiku 4.5 in practice?

Per token, yes. Per task, Vals found it costs more on every shared benchmark, because it reasons through far more steps. On Legal Research it took 59 steps versus 17 for Haiku 4.5. Turning down its effort setting is the main way to claw that back.

What changed for Sonnet 5.5?

Anthropic halved its cache reads to $0.10 per million tokens, with input at $2 and output at $10, which it estimates makes most agentic work about 20% cheaper. Claude Code limits did not change. More on Sonnet 5.5 in our launch coverage.

Should I use Haiku 5.5 for coding?

It is strong for a small model, with 90.4% on Vibe Code Bench and a big Code Arena gain, so it works well as a cheap sub-agent. For the hardest problems you still want a flagship. See our best AI models for coding rundown.

Sources

Further reading

Try all the models mentioned in this article

Admix gives you GPT-5, Claude, Gemini, and 350+ AI models in one app. Compare responses side by side. Free to start.

Start free on Admix

Related articles