
The Cheapest Way to Run a Coding Agent in 2026 Is to Stop Using One Model
Quick verdict
Frontier coding models are now close enough that the price you pay depends less on which single model you pick and more on how you split the work. Together says GLM-5.2 does about 80% of Sonnet 5's software engineering at roughly a fifth of the cost, and you can now plug it into Claude Code yourself. The people saving the most money are routing the cheap model to the bulk of the work and paying frontier prices only for the steps that need it. The people losing money are letting a provider route for them.
What actually shipped
Two things landed in the same week that make self-directed model routing practical for a solo developer, not just an enterprise platform team.
First, Together reported that GLM-5.2 reaches roughly 80% of Sonnet 5's software-engineering capability at about 20% of the price. That is a specific, testable claim about coding work, not a general leaderboard average.
Second, GLM-5.2 became selectable inside Claude Code through Hugging Face Inference Providers. That closes the gap between "there is a cheaper open model" and "I can actually use it in the harness I already run." You keep the Claude Code workflow and swap the engine underneath it.
The independent SWE-rebench numbers back up why this is tempting rather than reckless:
| Model | SWE-rebench solve rate | Tokens per task |
|---|---|---|
| Claude Opus 4.8 xhigh | 56.5% | 2.48M |
| GLM-5.2 | 51.1% | 2.62M |
| Gemini 3.5 Flash | 49.5% | 1.85M |
| MiniMax M3 | 45.6% | 6.89M |
| DeepSeek-V4 Pro | 42.7% | — |
GLM-5.2 lands five points behind a top-tier closed model on the same benchmark, at a fraction of the token price. For a lot of everyday coding, that is a trade worth making.
The three-model stack that costs a few dollars
The clearest template came from Mitchell Hashimoto, who described a planner, coder, judge workflow: a strong model plans the change, a second model writes the code, and a strong model judges the result. In his setup the plan and judge steps cost only a few dollars because they run on short, high-value prompts, while the expensive end-to-end loop, where one frontier model does everything from reading the repo to editing every file, is where the bill usually balloons.
The insight is that not every step needs your best model. Planning and reviewing are cheap to run and expensive to get wrong, so pay up there. The long middle, reading files, drafting edits, running tests, is where token counts pile up, and that is exactly where a model at a fifth of the price earns its keep. Simon Willison made a related point that the real bottleneck with coding agents is no longer raw code generation but understanding enough to participate in what the agent is doing. Splitting the work by step is one way to stay in the loop instead of handing a black box your credit card.
Why it matters
The counterexample showed up the same week. A user with "Claude Fable 5" selected watched a single session run up a $321.53 bill after Anthropic's guardrails quietly rerouted the work to Opus 4.8. Commenters called it the "Opus sandwich": you pick a cheaper tier, a classifier decides your request needs a bigger model, and you pay the bigger model's price plus a likely cache miss when the context moves across the boundary. The suggested fix in that thread was to set fallback off so the model you chose is the model you get.
That is the whole argument for doing your own routing. When a provider routes, it optimizes for its safety and capacity constraints, and the cost lands on you with no warning. When you route, you decide which step is worth frontier prices. A popular r/LocalLLaMA thread made the deeper point that closed products are often a whole pipeline, routing, hidden prompts, tool calls, wrapped around a base model, which is part of why the raw open weights look closer than the marketing suggests once you handle the orchestration yourself.
Anthropic did make access less painful the same week, raising and simplifying API rate limits. That helps, but it does not change the routing math. Hugging Face's Clement Delangue framed the broader shift as open models becoming the sovereignty layer for developers, and the cost data is what turns that from a slogan into a budgeting decision.
Video: testing a cheaper Claude alternative
A hands-on walkthrough of pointing a cheaper GLM model at a Claude-style coding workflow, including the setup and where the quality actually drops off.
How to actually split the work
A practical starting point that mirrors the stacks people are reporting:
- Plan with a frontier model. One short, expensive call that sets the approach. Cheap to run, costly to skip.
- Implement with a cheaper model. GLM-5.2 or a comparable open model handles the file-by-file grind where tokens accumulate.
- Judge with a frontier model. A second short call reviews the diff before you trust it.
- Turn provider fallback off where you can, so you are not silently upgraded to a pricier tier mid-task.
FAQ
Is GLM-5.2 actually good enough to replace a frontier model for coding?
For most day-to-day work, close. It sits about five points behind top closed models on SWE-rebench at a fraction of the price. For the hardest planning and review steps, keep a frontier model in the loop. See our GLM-5.2 breakdown for the catch nobody mentions.
Does routing to a cheaper model break my Claude Code setup?
No. GLM-5.2 is now selectable in Claude Code through Hugging Face Inference Providers, so you keep the harness and swap the engine. Our guide to the best AI coding agents covers which harness fits which workflow.
How do I avoid a surprise bill like the $321 Opus session?
Disable automatic fallback where the provider allows it, and route expensive steps deliberately instead of letting a classifier decide. The same discipline enterprises use to cut their AI spend by routing works for individuals.
Sources
- @togethercompute - GLM-5.2 hits ~80% of Sonnet 5 coding at ~20% of the price
- @zRdianjiao - GLM-5.2 now selectable in Claude Code via Hugging Face Inference Providers
- @mitchellh - planner, coder, judge stack costing only a few dollars
- @simonw - understand to participate, the antidote to cognitive debt
- @ClaudeDevs - Claude API rate limits raised and tiers simplified
- @ClementDelangue - open models as the sovereignty layer
- r/singularity - Fable session silently routed to Opus 4.8, $321 bill
- r/LocalLLaMA - the closed vs open gap may be orchestration, not the base model
Further reading
- GLM-5.2 is the first open model people actually call a daily driver
- Enterprises are cutting AI spend almost in half by routing, not by using less
- The best AI coding agents in 2026
- The best free AI coding agents (open source, BYOK)
- AI model routers explained: LLM routing and claude code router
- One subscription for all the AI models
Try all the models mentioned in this article
Admix gives you GPT-5, Claude, Gemini, and 350+ AI models in one app. Compare responses side by side. Free to start.
Start free on Admix