← All posts / Industry

Gartner's 'Inference Paradox': Agentic Workflow Costs to Grow Fivefold Through 2028

Gartner's August 17 report predicts AI inference costs per agentic workflow will rise more than 5x through 2028 as token efficiency gains are swallowed by increasingly complex, autonomous multi-model workflows — the 'Inference Paradox'.

Gartner's 'Inference Paradox': Agentic Workflow Costs to Grow Fivefold Through 2028

On August 17, 2026, Gartner published a prediction that cuts against the grain of nearly every AI cost headline this year: the cost of inference per agentic workflow will increase more than fivefold through 2028. Not per token. Not per model call. Per workflow — the full multi-step, tool-calling, self-questioning loop that defines an AI agent doing real work.

The claim lands at a strange moment. Just five months ago, in March 2026, the same firm predicted that by 2030, running inference on a one-trillion-parameter LLM would cost GenAI providers over 90% less than in 2025. Both forecasts can be true — and the tension between them is precisely the point of what Gartner now calls the Inference Paradox: better unit economics escalating the overall cost of AI, without providing a clear pathway to commensurate and predictable value.

Gartner identifies three fundamental dynamics driving token economics:

1. Foundational model cost economics are rapidly improving. Per-token prices keep falling, hardware keeps getting more efficient, and providers keep shipping cheaper tiers. This is the familiar half of the story — the one that has fueled a summer price war, with leading US model prices dropping nearly 25% since mid-July by some measures.

2. Efficiency unlocks more expensive models. Improved AI efficiency doesn’t just cut costs — it makes it economically viable to deploy larger, more powerful, and more expensive models for higher-value applications. Cheaper tokens don’t get banked as savings; they get spent on capability.

3. Sophisticated workflows consume far more tokens. An agentic workflow — where a model plans, calls tools, inspects results, revises, and iterates — burns orders of magnitude more tokens than a chatbot exchange. As products evolve from assistive features to multistep execution, the multiplier effect dominates.

Put together: tokens are becoming more cost-efficient, but not as quickly as AI capabilities — and the costs of those capabilities — are increasing. The rate of innovation is outpacing the cost curve.

Chatbot versus agent: a fivefold gap at minimum

Gartner analyst Will Sommer, Sr. Director Analyst, framed the gap in concrete operational terms. “Where a simple chatbot must read and interpret a query and quickly respond with a probabilistically reasonable answer, an AI agent must constantly reason, negotiate, and question itself.”

All of those responsibilities add up. Compared to a basic chatbot interaction, routing a task to an agentic reasoning model increases provider inference costs by at least five times — and often much more as task complexity grows.

His warning to product leaders is blunt: “Product leaders cannot rely on more efficient token economics to rationalize AI costs. Each successive generation of AI capability will necessitate more, and often more expensive, tokens. There is no reliable, economical one-size-fits-all model on the horizon. Producing competitive AI products will require developing and maintaining complex multimodel ecosystems.”

Why falling prices can mean rising bills

The Inference Paradox is essentially Jevons paradox applied to compute. When the price of a unit of intelligence falls, consumption doesn’t hold steady — it explodes. Each efficiency gain makes it rational to push AI into more steps of more workflows, at higher capability tiers, with longer reasoning chains. Enterprises that migrated workloads to cheaper models this summer and watched their total spend climb anyway have already lived through a private version of this dynamic.

Gartner’s June 2026 prediction that AI coding costs will surpass the average developer’s salary by 2028 was an early signal of the same phenomenon. Now the framing has generalized: it isn’t just coding agents, it’s the entire shift from assistive AI to autonomous execution that rewrites the cost structure.

The report also closes off an easy escape route. If no single model will be economical for everything, then the answer isn’t finding a cheaper default — it’s architecture. Ensuring ROI from advanced AI like reasoning agents demands, in Gartner’s words, “exponentially higher returns relative to basic models, or highly optimized inference-tiering, routing and orchestration to calibrate complex tasks relative to more cost-efficient intelligence.”

The strategic takeaway: tiering or bust

For product leaders, the actionable core of the report is inference tiering — routing each step of a workflow to the cheapest intelligence that can handle it. Simple subtasks go to small, fast, inexpensive models. Judgment calls, planning, and ambiguous reasoning go to frontier models. Orchestration layers decide which is which.

“Defaulting to generic autonomous intelligence will result in unbounded costs orders of magnitude higher than those of optimized product ecosystems,” Sommer said. The phrase “unbounded costs” is doing heavy lifting: without deliberate tiering, agentic systems can loop, retry, and over-reason their way through budgets with no natural stopping point.

This has practical consequences across the industry. Model providers will keep competing on sticker price per token, but the real competition is shifting to who can deliver the cheapest completed task — a metric that depends on routing intelligence, token efficiency, and workflow design as much as on list prices. Platform vendors that expose fine-grained model choice, caching, and prompt-context management are selling the picks and shovels of the tiering gold rush. And enterprises signing AI contracts based on per-token rates may want to renegotiate around per-task or per-outcome pricing before the fivefold curve catches up with them.

Context and caveats

Gartner has skin in this game — the full analysis is available to clients in a report titled The Inference Paradox: Inference Tiering Is Critical to Protect Margins, and the firm will elaborate on AI economics at its IT Symposium/Xpo events starting September 14 in Gold Coast, Australia. Predictions with a named framework and a conference circuit behind them should be read partly as agenda-setting.

But the underlying arithmetic is hard to argue with. Agentic workflows are structurally more expensive than chat, capability upgrades structurally cost more tokens, and adoption of agents is still in its early innings. Gartner’s own survey data from May 2026 found only 17% of enterprises had deployed AI agents, with another 42% planning to within twelve months. If even half of that pipeline materializes while per-workflow costs grow fivefold, the total AI inference bill of the industry grows by an order of magnitude — efficiency gains notwithstanding.

The uncomfortable conclusion: in AI economics, the unit price is falling, the unit is shrinking, and the bill is still going up. The winners of the next phase won’t be those who bet on any single model being cheap enough. They’ll be the ones who built the discipline — and the plumbing — to spend the least intelligence necessary at every step.