
AI Agents Pass Just 26% of Real Job Tasks, and 2.6% of the Hard Ones
Quick verdict
A new benchmark called Agents' Last Exam took the best AI agents available right now and gave them 1,500 real tasks pulled from 55 actual occupations. The headline result is sobering for anyone betting that agents replace knowledge workers this year: the top agent passed about 26% of tasks overall, and just 2.6% on the hardest tier. Every leading agent, Codex, Claude Code on Fable 5, and the open ALE Claw setup, landed within a couple of points of each other in the low 20s. Agents are not useless here. But the gap between a slick demo and a reliable hire is still wide.
What the benchmark does
Agents' Last Exam is co-led by UC Berkeley's RDI group, with more than 250 expert contributors across 100-plus institutions and support from Snorkel AI's open benchmarks program. The tasks are grounded in the O*NET occupational taxonomy the US government uses to define jobs, so they map to real work in 55 non-physical fields rather than puzzles invented for a leaderboard.
Agents get dropped onto real machines and have to produce a deliverable through a GUI or a command line, the same way a remote worker would. There are no human judges. Every task is code-graded and reproducible, so the agent is scored on the result, not on how convincingly it narrates its work. That removes the usual wiggle room where a model sounds competent without finishing the job.
The leaderboard, in plain numbers
Here is where the frontier agents landed as of the June 11 board. Pass rate is the share of tasks fully completed. The score column gives partial credit.
| Agent | Model | Pass rate | Score |
|---|---|---|---|
| Codex | gpt-5-5 | 24.0% | 42.8% |
| ALE Claw | gpt-5-5 | 23.0% | 45.8% |
| Claude Code | claude-fable-5 | 22.0% | 40.5% |
| OpenClaw | gpt-5-5 | 21.1% | 41.0% |
Two things jump out. First, the best general result tops out around 26%, which means roughly three out of four real tasks go unfinished. Second, on the "Last Exam" tier, the set of genuinely hard professional tasks, the average pass rate collapses to 2.6%. That is the number worth remembering. The hardest real work is still almost entirely out of reach.
Why it matters
Forecasts have been getting louder that agents will outperform humans at nearly every job by 2026 or 2027. This benchmark is a direct test of that claim under fair conditions, and the agents are nowhere close. The reason the demos look so much better than the scores is that a demo is a curated happy path. A real job is 1,500 messy tasks where finishing 22% of them gets you fired, not promoted.
For anyone planning to put agents into a workflow, the practical reading is to scope them tightly. They are strong on short, well-bounded tasks and weak on the long-horizon work that defines most actual roles. That matches what people are already seeing when they push AI sales agents or customer support agents past the easy first interactions. The same caution applies to coding: agents are useful copilots, but the benchmark is a reminder to keep a human reviewing the output rather than shipping it blind. If you are choosing tools, our notes on the best AI coding agents and the best models for coding hold up better when you treat them as assistants, not replacements.
The other quiet detail is how close the models are. Codex on GPT-5.5 and Claude Code on Fable 5 differ by two points. The capability gap between the top labs on real work is small, which is one more reason to keep access to several models instead of locking into one.
Video: can AI handle real work?
A grounded look at where agents help and where they stall on actual tasks, which lines up with what the benchmark measured.
FAQ
Does a 2.6% score mean AI agents are a dead end?
No. It means the hardest, long-horizon professional tasks are still unsolved. Agents already handle plenty of short, bounded work well. The benchmark sets a ceiling on what to trust them with unsupervised, not a verdict on the technology.
Why are all the top agents scoring so close together?
Because they run on a handful of frontier models, mostly GPT-5.5 and Fable 5, and the bottleneck is long-horizon reliability rather than raw model smarts. The harness around the model matters less than people assume once you grade on finished deliverables.
How is this different from coding benchmarks like SWE-Bench?
SWE-Bench measures one skill, fixing software issues. Agents' Last Exam spans 55 occupations and grades full job deliverables on real machines, so it captures breadth and long tasks instead of a single domain.
Should I still pay for an AI agent?
If your work is bounded and you review the output, yes. Given how close the models score, paying for access to several through one app that runs multiple AI models tends to beat betting on a single agent.
Sources
- @YiyouSun - announcing Agents' Last Exam and the headline results
- @SnorkelAI - on the benchmark design and contributor network
- Snorkel AI - Agents' Last Exam: can AI agents actually do real jobs?
- arXiv - Agents' Last Exam (paper)
- Agents' Last Exam - live leaderboard
Further reading
Try all the models mentioned in this article
Admix gives you GPT-5, Claude, Gemini, and 350+ AI models in one app. Compare responses side by side. Free to start.
Start free on Admix