
AT&T Routes 40% of Its AI to Open Models and Cut Coding Costs 56%
Quick verdict
AT&T just gave the clearest public number yet for the bet enterprises are making on open models. Roughly 40% of its employee AI usage already runs on open weights, the target is 60-70%, and the switch cut coding costs 56% for a 2% quality drop, all at a scale of 45 billion tokens a day. The lesson underneath the headline is the one Admix keeps making: you save money by sending each task to the cheapest model that can do it, not by using AI less.
What AT&T actually reported
The datapoint came through a widely shared summary from Hesam Sheikh of AT&T's internal deployment. The figures are specific. About 40% of employee AI usage already routes to open models, with a stated goal of pushing that to 60-70%. Coding costs are down 56%, and the quality cost of that move was 2%. The whole thing runs at 45 billion tokens per day, which is enough volume that a few percentage points of routing decide a real budget line.
Read those numbers together and the strategy is obvious. Frontier closed models stay in reserve for the hardest work. Good-enough open models take the broad middle of everyday requests, where the marginal quality difference does not justify the price gap. A 56% cost cut for a 2% quality dip is not a close call. It is the kind of trade any finance team signs off on without a meeting.
Why this one lands harder than the usual routing story
Plenty of companies have said they are cutting AI spend. What makes AT&T different is that it is a named, conservative, regulated enterprise putting a hard percentage on open-model adoption, not a startup chasing a cheaper bill. When a carrier says it wants 60-70% of AI usage on open weights, the "open models are not ready for serious work" argument gets harder to hold.
Amir Efrati framed it as a warning sign for the enterprise moat that OpenAI and Anthropic are counting on. The reasoning is straightforward: if the frontier labs only get paid for the hardest slice of demand, and open models eat everything else, the revenue pool per customer shrinks even as usage grows. Ollama, which hosts open models for exactly this kind of deployment, welcomed AT&T to the club, which tells you where the supply side thinks this is going.
The catch buried in the numbers
A 2% quality drop is an average, and averages hide the tasks where a cheaper model quietly gets it wrong. AT&T can absorb that because it has the engineering to route well: hard prompts go to the strong model, everything else goes to the cheap one, and someone watches the error rate. The savings come from the routing being good, not from open models being free. Skip the routing discipline and you get the quality drop without the full cost win.
There is also a supply-constraint subplot in the same week's news. Users on the closed side reported hitting usage caps on expensive plans, with some noting a $200 Pro plan could be burned through in a single heavy coding day. Open weights you host, or rent with predictable pricing, sidestep that ceiling. At that point an open option in your stack stops being only a cost play and becomes a continuity plan for when the premium endpoint throttles you.
Video: how model routing cuts AI costs
A short primer on the routing idea behind AT&T's savings, sending each request to the cheapest model that can handle it.
What it means if you are not AT&T
The enterprise playbook is personal budgeting with more zeros. If you pay for ChatGPT Plus, Claude Pro, and Gemini Advanced as three separate subscriptions, you are running the opposite of what AT&T is doing. Three flat fees, each locked to one model, with no way to push the easy questions toward the cheaper option. AT&T's version of "cheaper defaults plus routing" at the personal scale is a single account that reaches every model and lets you pick per task. That is the whole case for an AI aggregator over a stack of subscriptions, and it is the same math AT&T ran at 45 billion tokens a day.
For coding specifically, the AT&T number tracks what teams already see: a mid-tier or open model handles most edits fine, and the flagship earns its price only on the genuinely hard multi-file work. Our breakdown of cutting coding-agent costs with routing walks through how to set that up, and the best AI coding agents covers which tools route for you.
FAQ
Are open models actually as good as GPT-5 or Claude?
Not on the hardest tasks, and AT&T's own numbers admit a 2% quality gap. The point is that most everyday requests are not the hardest tasks, so a cheaper model clears the bar at a fraction of the cost. You keep the frontier model for the work that needs it.
Can an individual really route like an enterprise?
Yes, without building anything. An aggregator that gives you access to multiple models under one plan is the consumer form of routing. You pick the cheap model for simple prompts and switch to the flagship when a task is hard. See saving money on AI subscriptions.
Does routing to cheaper models put my data at risk?
It depends on the host. AT&T runs open models on infrastructure it controls, and hosted open-model providers increasingly offer zero data retention. That is often a privacy improvement over sending everything to one closed vendor, not a downgrade.
Sources
- @Hesamation - AT&T's internal AI deployment: 40% open now, 60-70% target, coding costs down 56% at 45B tokens/day
- @amir - why AT&T's shift is a warning sign for the OpenAI/Anthropic enterprise moat
- @ollama - welcoming AT&T to open models
Further reading
Try all the models mentioned in this article
Admix gives you GPT-5, Claude, Gemini, and 350+ AI models in one app. Compare responses side by side. Free to start.
Start free on Admix