AT&T Cut Its AI Bill 56% With a Router: The Quiet Rise of Enterprise 'Tokenomics'
By routing easy queries to Llama and Gemma, AT&T cut some coding-task costs 56% at a 2% quality cost — and now wants 60-70% of its 45 billion daily tokens on open models.
The most consequential enterprise AI story of the past week did not involve a new model, a new chip, or a new benchmark. It involved a 140-year-old telephone company deciding which existing model should answer which question — and publishing the arithmetic.
On August 20, The Information reported that AT&T has cut the cost of coding and other advanced AI tasks by as much as 56% by routing employees’ queries to cheaper models when the job does not require a frontier one. The quality cost, according to AT&T’s own measurements: roughly 2%.
The numbers behind the numbers are even more striking. AT&T now processes an average of 45 billion AI tokens per day across its workforce — up from about 8 billion a day a year ago, a more than fivefold increase driven by the shift from occasional chatbot queries to always-on coding agents and document workflows. Roughly 40% of employee queries already run on open-source or open-weight models. Mark Austin, the AT&T vice president who oversees employee-facing AI, says the company wants that share to reach 60% to 70% in the coming years.
This is not a pilot. It is a Fortune 500 company re-architecting how it buys intelligence — and it is the clearest signal yet that the enterprise AI market is entering the phase analysts have started calling tokenomics: the discipline of managing token spend the way CIOs once managed cloud bills.
What AT&T actually built
The centerpiece is deceptively simple: a model router. AT&T uses LiteLLM, an open-source proxy that sits between employees and every model the company has access to. When a request arrives, the router classifies its complexity — using heuristics, keyword rules, and an LLM classifier — and decides whether it genuinely needs Claude or GPT-class capability, or whether Llama, Gemma, or Nvidia’s Nemotron can handle it.
- Simple tasks — document summaries, code explanations, formatting, routine lookups — go to open-weight models like Meta’s Llama and Google’s Gemma, which AT&T serves itself at commodity cost.
- Demanding tasks — complex code generation, nuanced analysis, high-stakes reasoning — still go to Anthropic and OpenAI frontier models, where the premium buys real capability.
- The router is the policy. Cost control is no longer a matter of telling employees which tools to use; it is an infrastructure decision made per-request, invisible to the user.
AT&T’s stated goal is not to eliminate frontier-model spending but to hold it flat. As usage compounds — remember, 8 billion to 45 billion daily tokens in twelve months — keeping the Anthropic and OpenAI line item flat while total consumption grows fivefold means the marginal token must increasingly be served by open models. That is the whole strategy in one sentence.
Notably, AT&T is drawing one line: it is not using Chinese open-weight models from DeepSeek or Moonshot, though Austin says the company is actively evaluating the risk picture. The geopolitical dimension of open weights is now a procurement question, not just a policy debate.
The 56% number, examined
Enterprises and vendors throw cost-saving claims around constantly; this one deserves a closer look before it becomes folklore.
First, the caveats. The 56% figure covers some coding tasks — not all AI spending, and not a controlled benchmark. The “2% decline in quality” is AT&T’s own assessment, reported secondhand, with no published methodology. A two-point quality drop on an internal evaluation is not the same as a two-point drop on every task an employee might attempt; tasks where open models fail outright may simply be excluded from the measured set.
Second, the direction of the result is corroborated elsewhere. Nvidia reported a 58% cost reduction in its own SWE-Bench evaluation when its NeMo Switchyard router sent routine agent steps to the smaller Nemotron 3.5 Lightning and reserved a frontier model for genuinely hard reasoning. Ramp’s expense data shows enterprises actively repricing their model mix. And Goldman Sachs’ equity research — more on that below — has built a thesis on exactly this behavior spreading.
Third, the mechanism is credible even if the headline is fuzzy. Most agentic workloads are dominated by high-volume, low-difficulty calls: tool-calling loops, summarization, reformatting, retrieval synthesis. If 80% of calls in a coding agent’s loop don’t need frontier capability, routing them away is arithmetic, not innovation. The 56% is best read as the measured savings on the subset of tasks where routing applied — which is still an enormous number for a company burning 45 billion tokens a day.
Why open weights are suddenly good enough
Austin’s framing of the capability gap is the most quotable detail in the reporting: AT&T finds open-source models generally six to ten months behind frontier models, but the gap is narrowing — and today’s open models are “just as good or better” than older paid models from Anthropic and OpenAI.
That last clause is the quiet killer. It reframes the buying decision entirely. If your alternative to Llama this year is last year’s Claude at this year’s price, the open model wins on both cost and capability. Frontier labs have effectively been competing against their own back catalog — and the back catalog keeps getting open-sourced by someone else.
Three forces converged to make this workable in 2026:
- Open models crossed the usefulness threshold. Qwen, Llama, Gemma, and Nemotron families now handle the long tail of enterprise tasks — summarization, classification, code explanation — at quality levels employees cannot distinguish from frontier output.
- Routing infrastructure matured. LiteLLM, OpenRouter (just acquired by Stripe for over $7 billion), Nvidia’s Switchyard, and every cloud provider’s built-in routing make per-request model selection a configuration file, not a research project.
- Token billing made costs visible. The industry’s shift from flat subscriptions to metered, token-based billing turned AI spend into a line item CFOs can see — and once they can see it, they manage it. The two-year era of “tokenmaxxing,” where employees were pushed toward the biggest models with no eye on cost, is ending.
The Goldman counter-thesis: cheap tokens may rescue Big Tech
Here is where the story inverts. The intuitive read is that open models threaten the AI establishment: if Llama is good enough, why pay Anthropic?
Jim Covello, head of Goldman Sachs equity research, argues the opposite may hold for infrastructure owners. Cheaper open models make it more likely that enterprises can implement AI profitably — and profitable use cases mean more queries, more tokens, and more demand for the data centers and GPUs the hyperscalers are frantically building. “It’s really good for the hyperscalers,” Covello said, “because it’s more likely that you’re going to be able to profitably fill up all this capacity that you’re adding.”
The market context matters here. Big Tech’s AI capital expenditures have become one of Wall Street’s central anxieties: for years, every capex announcement was rewarded, but over the last several quarters investors have started asking the harder question — not whether the capacity can be built, but whether profitable demand exists to fill it. AT&T-style tokenomics cuts both ways. It pressures premium model pricing while simultaneously expanding the universe of economically viable AI applications. The net effect on any given company depends on which side of that ledger it sits.
For the model labs themselves, the pressure is more direct. If AT&T’s playbook — hold frontier spending flat, route the rest to open weights — becomes standard practice across the Fortune 500, the labs’ revenue growth has to come from capability differentiation, not volume. That is a much harder business.
What it means for everyone else
AT&T is not an AI sophisticate; it is a cost-disciplined telecom with 100,000+ employees and ordinary enterprise workloads. That is precisely why its numbers matter. If routing works there, the burden of proof shifts to every enterprise still sending every query to the most expensive model by default.
The practical checklist falls out of AT&T’s experience:
- Measure token flows before optimizing them. AT&T’s tokenomics program started with visibility — knowing that usage grew 8B → 40B → 45B daily tokens is what made the problem tractable.
- Classify your workload mix. Most enterprises discover the same distribution: a large majority of requests are simple, a minority are hard. Routing exploits that asymmetry.
- Keep an escape hatch to frontier capability. The goal is spending flat, not zero — developers keep Claude and GPT access for work that justifies it.
- Treat model choice as plumbing, not identity. The interesting architectural decision in 2026 is not “which model did you buy” but “what is the cheapest model that reliably handles this job.”
The hyperscale AI story of the last two years was about capability: bigger models, faster chips, longer contexts. The AT&T story suggests the next phase is about economics: the same models, routed intelligently, at a fraction of the price. Not one frontier lab shipped a smarter model in the week this news broke — and the enterprise cost curve bent further than it had all month.
The engine is commodity. The chassis is where the money is.
Sources
- [1] https://www.theinformation.com/newsletters/applied-ai/t-using-open-source-models-curb-anthropic-bills
- [2] https://www.pymnts.com/news/artificial-intelligence/2026/att-slashes-ai-costs-by-adopting-model-routers-and-open-source/
- [3] https://techstartups.com/2026/08/21/goldman-sachs-says-open-source-ai-could-be-big-techs-unexpected-winner-as-att-cuts-model-costs-by-56/
- [4] https://www.fierce-network.com/cloud/open-models-are-driving-atts-ai-tokenomics-strategy
- [5] https://www.wsj.com/cio-journal/why-at-t-is-betting-big-on-open-weight-ai-a0ea03b1
- [6] https://www.tmforum.org/news-insight/newsroom/tokenomics-the-ai-cost-challenge-telcos-can-t-afford-to-ignore