Enterprises Are Cutting AI Spend Almost in Half by Routing, Not by Using Less

Enterprises Are Cutting AI Spend Almost in Half by Routing, Not by Using Less

6 min readJune 30, 2026

Quick verdict

Companies are not abandoning AI to save money. They are getting smarter about which model answers which request. A widely shared UBS summary says 60% of firms curbing AI spend are shifting to cheaper and open models while reserving premium ones for hard tasks. Coinbase says the same playbook, cheaper defaults plus automated routing plus caching, cut its AI bill nearly in half even as token usage went up. The takeaway for anyone paying for AI: you save by routing, not by rationing.

What the numbers actually say

The headline came from a UBS report summarized by Rohan Paul: among companies pulling back on AI spend, roughly 60% are moving to cheaper and open-source Chinese models, and many are using model routing to keep premium models in reserve for the genuinely difficult work. That reframes the "AI spend is out of control" story. The fix most teams reach for is not less AI. It is spending the same budget across a wider set of models so the expensive ones stop carrying easy requests.

Coinbase put real numbers on it. CEO Brian Armstrong described an internal playbook built on five moves: cheaper default models, automated routing, cache-aware requests, leaner context windows, and better visibility into where the spend goes. The result was an AI bill that dropped close to half even as the company's token usage grew. Usage up, cost down, because most of that usage stopped hitting the priciest endpoint.

Hugging Face's Clement Delangue made the supply-side version of the argument: a large share of workloads could run locally or on cheaper specialized models if routing were easier to set up. The bottleneck is not capability. It is the plumbing that decides where each request goes.

The five levers, in plain terms

  • Cheaper defaults. Make a small or mid-tier model the first responder. Most prompts are simple, and a flagship answer is wasted money on them.
  • Automated routing. Send the hard prompts to the strong model and everything else to the cheap one, based on the task rather than a coin flip.
  • Cache-aware requests. Reuse the parts of a prompt that do not change. LangChain's Harrison Chase noted that Manus treats KV-cache hit rate as maybe the single most important metric for a mature agent, because re-paying for the same context on every call is where budgets quietly bleed out.
  • Leaner context. Stop stuffing the whole history into every call. Shorter inputs cost less and often answer better.
  • Visibility. You cannot cut what you cannot see. Per-request cost tracking is what turns "AI is expensive" into a list of specific things to fix.

None of this is exotic. It is the same discipline that infrastructure teams have always applied to compute bills, now pointed at tokens. Vendors are leaning in too: Baseten showed live draft-model training for speculative decoding with a 20% bump in median acceptance rate, and Google Research published a way to retrofit multi-token prediction onto frozen models for on-device speedups. The whole stack is shifting toward serving the same quality for fewer dollars.

Why open weights keep coming up

Routing and open models travel together for a reason. Once you accept that a cheap model can handle most requests, an open-weight model you run yourself becomes a serious option for the bulk of the load. It is the "own versus rent" calculation: rent the frontier for the few prompts that need it, own a capable open model for the rest. The recent momentum behind open releases like GLM-5.2 is part of this, and so is the steady stream of strong open checkpoints we covered in our guide to open-source AI models.

There is a risk angle too. When closed-model access can be gated or delayed, as it was during the restricted GPT-5.6 preview, an open option in your stack stops being a cost play and starts being a continuity plan.

Why it matters if you are not a company

The enterprise playbook is just personal budgeting with a bigger invoice. If you pay for ChatGPT Plus, Claude Pro, and Gemini Advanced separately, you are running the opposite of routing: three flat fees, each locked to one model, with no way to send the easy questions to the cheap option. The individual version of "cheaper defaults plus routing" is one account that reaches every model and lets you pick per task. That is the entire case for an AI aggregator instead of stacking subscriptions, and it is the same logic Coinbase used at scale. For the wider breakdown of where subscription money leaks, see our notes on saving money on AI subscriptions.

Video: model routing and AI cost control

A walkthrough of how routing between models cuts AI spend without dropping quality.

FAQ

Does switching to cheaper models hurt quality?

Not if you route. The point of routing is that a small model handles the easy majority of requests while the flagship still answers the hard ones. Quality drops only when you force one model to do everything, which is exactly what flat single-model subscriptions push you toward.

What is the biggest single cost lever?

For agents and repeated workflows, caching. Re-sending the same unchanged context on every call is where budgets bleed, which is why teams track KV-cache hit rate so closely. For one-off chats, choosing a cheaper default model matters most.

Are open models really good enough for production?

For a large share of everyday tasks, yes, which is why UBS found so many cost-cutting firms moving to them. The frontier still wins on the hardest problems, so the common setup is open models for volume and a premium model held in reserve. See our open-source AI models guide for where they stand today.

How do I apply this as an individual?

Stop paying separate flat fees for each model. Use one account that reaches all of them and pick the cheapest model that can do each job. That is routing at a personal scale, and it is the core argument for an AI aggregator.

Sources

Further reading

Try all the models mentioned in this article

Admix gives you GPT-5, Claude, Gemini, and 350+ AI models in one app. Compare responses side by side. Free to start.

Start free on Admix

Related articles