Claude Opus 5 Launched, and the Benchmark Fight Is Already Louder Than the Model

Claude Opus 5 Launched, and the Benchmark Fight Is Already Louder Than the Model

6 min readJuly 26, 2026

Quick verdict

Anthropic launched Claude Opus 5, and the first day was less about a spec sheet and more about a fight over how to measure it. Epoch put it at 159 on its Capabilities Index, one point under Fable 5, but level with Fable 5 at 161 on the software engineering track. People who ran it on real coding and browser tasks came away saying it feels much stronger than a one-point gap suggests. If you buy AI by the task, the takeaway is simple: Opus 5 is a top-tier coding and agent model, and the public benchmarks have not caught up to how it behaves in practice yet.

What actually shipped

Opus 5 arrived as a frontier model aimed squarely at coding and agent work, and the early signal came from evaluations and hands-on demos rather than a big launch table. The concrete numbers came from Epoch:

  • Claude Opus 5 scored an Epoch Capabilities Index (ECI) of 159, which Epoch described as slightly below Fable 5's 161
  • On software engineering, Opus 5 hit SWE-ECI of 161, matching Fable 5
  • Community reaction pegged the overall score at only about 1 point above Opus 4.8
  • Nous Research added Opus 5 to its portal with a 20% discount across models
  • Arena said first impressions were live and real-world leaderboard rankings were still coming
MetricClaude Opus 5Fable 5
ECI (overall)159161
SWE-ECI (software engineering)161161

The split in those two rows is the story. Opus 5 trails on the broad omnibus score but ties on the coding-specific one, which lines up with what Claude models have been known for.

Why the benchmark and the vibes disagree

The ECI number landed and people pushed back fast. One widely shared response called Opus 5 "incredibly underrated," pointing out that a single point over Opus 4.8 felt far too small for a model that seemed better at nearly everything in day-to-day use. The same account used that gap to argue for harder public benchmarks, on the theory that the current ones are saturating and can no longer separate frontier models cleanly.

There was also a genuinely odd result. One evaluator found Opus 5 scoring better on FrontierCode at medium effort than at high effort, even though more effort helped on other tasks. More thinking is supposed to mean better answers, so a model that gets worse when you let it think harder points to either task-specific search tradeoffs or plain evaluation noise. Either way, it is a reminder that a single index number hides a lot of behavior underneath.

Where it looked strong: coding and browser control

The praise that carried weight came from people running actual work. Microsoft's Kevin Scott, posting as Mikhail Parakhin, said "best-of-n rules" and reported a clear head-to-head win over Fable for math and coding, adding that he wished it were available in Codex. On the agent side, one developer said Opus 5 opened a browser and canceled a ChatGPT Pro subscription on its own, then followed up with "this thing can really drive a browser." Those are demos, not systematic evals, but they map to the category that public benchmarks miss the most: an agent that can drive a real interface to finish a real task.

That is the wedge. Static question-and-answer benchmarks compress coding, tool use, latency, and best-of-n gains into one score. The behaviors people are excited about here, browser automation and long agentic loops, are exactly the ones that a single ECI figure flattens.

Why it matters

Opus 5 is not landing in a quiet field. It is being judged next to Fable 5, GPT-5.6, Grok 4.5, and open-weight releases like Kimi K3, where coding ability is the main dividing line and cost is the next one. In that context, the two-point overall gap matters less than one detail: even the people knocking the benchmark are arguing about how much better Opus 5 is, not whether it belongs at the frontier. When the debate is over the size of the win, the model has already made its case. For teams choosing where to spend, the practical read is that Anthropic held its coding lead and that the leaderboards will likely drift toward the hands-on verdict over the next few weeks. See our OpenAI vs Anthropic vs Google breakdown for how the three stack up right now.

Video: testing Opus 5 on real work

A grounded, non-hype walkthrough of where Opus 5 shines and where it stumbles, which is a useful counterweight to launch-day enthusiasm.

FAQ

Is Claude Opus 5 better than Fable 5?

On Epoch's overall index, no. Opus 5 scored 159 against Fable 5's 161. On the software engineering track they tied at 161, and several users who ran their own coding tests said Opus 5 came out ahead in practice, especially with best-of-n sampling. It depends on whether you weigh the omnibus number or the coding-specific one.

How much better is Opus 5 than Opus 4.8?

By the headline ECI, only about a point, which is what sparked the "underrated" complaints. Users report the real-world jump feels larger than that gap, particularly on browser control and agentic tool use, but there is no independent leaderboard confirming that yet. For the prior generation, see our Claude Opus review.

Can Opus 5 actually control a browser?

Early demos say yes. One developer reported it opened a browser and canceled a subscription without hand-holding. Those are isolated examples rather than a measured success rate, so treat them as promising rather than proven.

Where can I use Claude Opus 5?

Through Anthropic directly, and through third-party portals like Nous Research, which added it with a 20% discount at launch. If you are comparing coding tools around it, our guide to the best AI coding agents covers where each one fits.

Sources

Further reading

Try all the models mentioned in this article

Admix gives you GPT-5, Claude, Gemini, and 350+ AI models in one app. Compare responses side by side. Free to start.

Start free on Admix

Related articles