Claude Opened Its 1,000-Agent Workflows to Everyone. The First Cost Numbers Say Go Slow.

Claude Opened Its 1,000-Agent Workflows to Everyone. The First Cost Numbers Say Go Slow.

6 min readOctober 11, 2026

Quick verdict

Anthropic moved its multi-agent features out of the waitlist this week. Claude Managed Agents can now split one request into a phased plan and hand it to as many as 1,000 agents in a single run, and every Pro and Max user who was waiting for Claude Code Projects got let in. The capability is real and some of the demos are striking. The money math is the part most people will skip and then regret. Anthropic itself tells you to start small because token use runs high, and the first independent test of agent teams found they cost roughly two to five times more than a single model while almost never producing better code. Turn it on for the jobs that genuinely split into parallel pieces. For everything else, one good model is still cheaper and usually just as good.

What actually shipped

Two things went live at once. Claude Managed Agents dynamic workflows entered public beta: a lead agent writes a phased plan, fans the work out to the sub-agents, and merges what comes back. You switch it on with the multiagent_20261001 flag, and the cap is up to 1,000 agents per run. Alongside it, Claude Code Projects cleared its waitlist, so all the Pro and Max users who signed up now have it. Each project runs its tasks as parallel threads, and sessions can run locally rather than only in the cloud.

The warning is in Anthropic's own rollout notes, not in a critic's thread. The guidance says to begin with scoped tasks because token consumption can be high. That is the company telling you the headline number, 1,000 agents, is a ceiling and not a setting you want to reach by default. One more pricing detail landed the same day and points in the same direction: Opus 5.5 fast mode is out, but it bills against usage credits and is not part of a subscription, so the fast path is a metered path.

The benchmark nobody running a swarm wants to read

Vals AI ran the obvious experiment. It put GPT-6 Sol and Claude Opus 5.5 through Vibe Code Bench twice, once as a single model and once as a team of agents, and compared both cost and quality. The teams were far more expensive and the extra spend mostly bought nothing.

Setup Cost vs single model Quality change
Agent teams, overall 1.8x to 5.1x more No significant gain in most runs
GPT-6 Sol, medium effort Higher +7.3 points (the one clear win)
Opus 5.5, max effort Highest No significant gain

The behavior behind the numbers is the interesting part. Sol delegated in parallel along architectural lines, which is why its medium-effort run was the single case that improved by a meaningful margin. Opus ran sequential waves instead, reaching about 6.8 sub-agents and roughly 1,140 sub-agent tool calls per app at max effort, and all that activity did not move the score. More agents turned into more tool calls, more tokens, and a bigger invoice, without better output.

Why it matters

Swarms are not a scam. They are a tool with a narrow fit, and the fit is tasks that actually decompose into independent parallel work. Prime Intellect's Prime Agent is the clearest example from this week: it rewrote itself in Rust over two weeks using a swarm of more than 2,000 agents, 10,000-plus sandboxes, and over 200 billion GLM-5.3 tokens, and the result now reaches usable input about 13 times faster and uses 83 percent less startup memory. That is a job where parallelism pays, and the token bill was the point of the exercise. Devin is leaning the same way, letting one agent spawn a tree of managed Devins so that wall-clock time tracks the slowest branch rather than the sum of every step.

For the daily work most people do with a coding agent, the Vals AI result is the one to keep in mind. If you pay by the token, every sub-agent you spin up is another meter running. A single capable model finishing a task in one pass is cheaper than a team of agents arriving at the same answer after a thousand tool calls. The question to ask before you turn the swarm on is whether the task splits cleanly, not whether the feature is available. Most of the time the honest answer is no, and the cheapest routing decision is to not route at all. For the cases that do split, scope them tightly, watch the credit meter, and compare the bill against a plain single-model run before you make it a habit.

Video: how Claude's multi-agent workflows work

A walkthrough of the orchestration pattern, where a lead agent plans and sub-agents execute, and where the token cost comes from.

FAQ

What are Claude Managed Agents?

A dynamic-workflow feature, now in public beta, where a lead agent breaks a request into a phased plan, hands the pieces to sub-agents, and merges the results. You enable it with the multiagent_20261001 flag, and a single run can fan out to as many as 1,000 agents.

Do agent teams write better code than a single model?

Usually not. Vals AI's test on Vibe Code Bench found teams cost 1.8 to 5.1 times more with no significant quality gain in most runs. The one clear win was GPT-6 Sol at medium effort, which improved by 7.3 points. If you want a single-model baseline to compare against, start with our roundup of the best AI coding agents.

How do I keep multi-agent costs under control?

Scope each run tightly, since Anthropic warns token use is high, and remember that Opus 5.5 fast mode bills against usage credits rather than your subscription. For the broader approach, see our guide to cutting coding-agent costs with model routing.

Is this the same as Cursor Projects?

The idea is similar. Both put a coordinator in charge of a fleet of sub-agents. We covered Cursor's version in Cursor Projects puts a coordinator agent in charge of a fleet of coding subagents.

Sources

Further reading

Try all the models mentioned in this article

Admix gives you GPT-5, Claude, Gemini, and 350+ AI models in one app. Compare responses side by side. Free to start.

Start free on Admix

Related articles