The Best AI Finished Only 3% of Real Work Tasks, and the Winner Cost $31 Each

The Best AI Finished Only 3% of Real Work Tasks, and the Winner Cost $31 Each

5 min readJune 20, 2026

Quick verdict

Artificial Analysis launched a benchmark called AA-Briefcase that does something most evals avoid: it tests AI on long, messy, multi-week projects instead of one-shot questions. Two numbers stand out. The top model satisfied every rubric criterion on only 3% of tasks, so frontier AI is nowhere near "set it and forget it" on real knowledge work. And the cost spread is enormous: the winner, Claude Fable 5, averaged $31 per task, while GLM-5.2 did respectable work for $2.40. If you are paying for AI to do actual deliverables, the gap between "best" and "good enough" is where all the money lives.

What AA-Briefcase actually tests

Most benchmarks reward a model for a clean answer to a self-contained prompt. AA-Briefcase, introduced by Artificial Analysis, does the opposite. Each task is built around a multi-week project with thousands of fragmented inputs: Slack threads, email chains, scattered documents. The model has to dig through the corpus and produce a real deliverable, like a financial model or a board deck, that gets graded against a detailed rubric.

That design surfaces two things at once: quality and economics. It measures whether the output is actually correct and complete, and it tracks what the run costs, because long-horizon tasks burn a lot of tokens. A model can look brilliant on a trivia leaderboard and still be too expensive or too unreliable to hand a week-long project.

The scores, and the price of each

Here is how the headline models landed on the new eval.

ModelScore (Elo)Avg cost per task
Claude Fable 51587$31.00
Claude Opus 4.81356$10.40
GPT-5.5 (xhigh)not the leader$3.68
GLM-5.21266$2.40

Fable 5 won, but it cost three times what Opus 4.8 did and roughly thirteen times what GLM-5.2 did. Opus sits in the middle on both axes. The open-weight GLM-5.2 scored about 80% of the leader's Elo at under a tenth of the price. None of this is free quality: Artificial Analysis placed GLM-5.2 between GPT-5.5 and Opus 4.8 on the same eval, which is a strong showing for a model you can download.

Why the 3% number matters more than the ranking

The leaderboard order is the part everyone screenshots. The part worth keeping is that even the best model fully satisfied the rubric on just 3% of tasks. Real, fragmented, multi-week knowledge work is still hard for AI, and most "the agent did my whole job" demos are running on tasks far simpler than what AA-Briefcase throws at them.

For anyone paying for these tools, that reframes the buying question. You are not choosing the model that finishes the job, because none of them reliably do that yet. You are choosing how much to spend on a strong first draft that a human still has to finish. At $31 a task, you want to be sure the extra quality over a $2.40 run is worth it. Often it will not be.

What this means if you pay for AI

The cost spread is the practical takeaway. Most people on a flat $20-a-month plan never see per-task pricing, but it is exactly what you are exposed to the moment you run an agent on real work through an API. Routing every task to the most expensive model is how bills explode for marginal gains. The smarter pattern is to send routine work to a cheaper model like GLM-5.2 or GPT-5.5 and save the premium model for the cases where the rubric difference actually pays for itself. That kind of model routing is the whole argument for not locking yourself into a single provider's subscription, which we get into in our guide to what AI subscriptions really cost.

Video: how AI agents are evaluated

Context on why long-horizon agent benchmarks like AA-Briefcase look so different from the leaderboards you usually see.

FAQ

Does a higher Elo mean I should pay for the top model?

Not automatically. Fable 5 led but cost about thirteen times what GLM-5.2 did for a roughly 25% Elo edge. Whether that is worth it depends on how much the last bit of quality is worth on your task. For most routine work it is not.

Why did even the best model finish only 3% of tasks completely?

Because AA-Briefcase grades multi-week projects with thousands of fragmented inputs against strict rubrics, not tidy single prompts. Real knowledge work is messy, and current models still produce strong drafts rather than finished deliverables. Our writeup on the agent benchmark that mirrors real jobs covers the same pattern.

How can I avoid overpaying across models?

Match the model to the task instead of defaulting to the priciest one, and use a setup that lets you switch models freely. See how one subscription can cover multiple models for the routing approach.

Sources

Further reading

Try all the models mentioned in this article

Admix gives you GPT-5, Claude, Gemini, and 350+ AI models in one app. Compare responses side by side. Free to start.

Start free on Admix

Related articles