The CPU Comeback: Agentic AI's Unexpected Bottleneck Is the Humble Processor
IEEE Spectrum reports that agentic AI has flipped CPUs into the new performance bottleneck — Intel server chips are sold out, AMD doubled forecasts, and AWS is rationing CPU cycles.
The AI infrastructure story has, for three years, been a GPU story. Nvidia’s accelerators became the most sought-after silicon on Earth, memory followed, and CPUs — the general-purpose workhorse at the heart of every server — were treated as an afterthought, a commodity boxed alongside the parts that actually mattered.
That story is now officially over. In an IEEE Spectrum feature published August 16, 2026, contributing editor Matthew S. Smith lays out the case that agentic AI has inverted the data center’s performance hierarchy: the CPU, long the boring sibling of the GPU, has become the new bottleneck for AI workloads.
How Agents Broke the CPU Story
The shift caught even the largest cloud operator off guard. Earlier this year, according to the report, Amazon Web Services delivered a blunt mandate to its engineers: conserve CPU cycles at all costs. AWS has reportedly experienced an explosion in wait times for CPU server capacity as AI workloads strain its cloud infrastructure.
The mechanics are straightforward once you look past the GPU-centric framing. In the chatbot era, an LLM call was mostly a one-shot inference: prompt in, tokens out, done. CPUs handled setup and teardown but did little heavy lifting. Agentic systems work differently. A model that operates autonomously — spawning sub-agents, making API calls, reading and writing files, downloading packages, running code — pushes an enormous amount of orchestration work onto general-purpose compute.
Matt Kimball, vice president and principal datacenter analyst at Moor Insights & Strategy, puts the multiplication problem in stark terms: an enterprise workload that spawns 100 agents becomes “tens of thousands, hundreds of thousands, or millions of agents” once rolled out across a company, with agents calling tools and talking to each other through protocols like Anthropic’s model context protocol.
Seven of Eight Pipeline Stages
The industry’s own numbers back the anecdote. Souvik Kundu, senior staff research scientist at Intel, notes that many components of an agentic AI task are inherently CPU-based jobs: parsing model output, deciding which tool to invoke, making the API call or running the code, collecting the result, and feeding it back into the model.
AMD’s framing is even sharper. Madhu Rangarajan, vice president of compute and enterprise AI at AMD, says that in the company’s testing, seven of the eight stages in realistic agentic AI pipelines run entirely on the CPU.
The result is a stutter-step execution pattern that wastes both halves of the machine. Kundu, working with researchers from the Georgia Institute of Technology, found that the CPU is often idle while LLM inference runs on the GPU — and, conversely, the GPU is often idle when tool calls execute on the CPU. Their proposed scheduling optimizations can cut end-to-end latency by up to 1.8x under sustained load, but that gain chases a moving target: agentic systems generate work at machine speed and multiply it as they go.
Tokenization: The Hidden Tax
The Georgia Tech research also surfaced a subtler culprit: tokenization, the process that converts text into the integer token IDs a model can digest. Unlike the massively parallel matrix math that dominates LLM inference, tokenization is branchy, data-dependent, sequential string manipulation — classic CPU territory.
Euijun Chung, a PhD student at Georgia Tech who co-authored the complementary paper, explains why agents make this worse. A model that makes a tool call must parse and tokenize the result — and with long contexts, that means re-tokenizing the entire ongoing sequence each time. “If you have an ongoing sequence of, say, 100,000 tokens, and you have a tool result of 1,000 tokens, the tokenizer will have to tokenize the whole sequence again,” Chung says. “And you have to do tokenization at every agentic tool call.”
The latency consequences compound with sequence length. In test runs at longer sequence lengths, adding CPU cores reduced time-to-first-token latency by roughly 1.5x to 7x. Chung’s team tested smaller models — Alibaba’s Qwen 3-30B and Meta’s Llama 3.1-70B — and speculates larger models may face less dramatic relative bottlenecks, but he expects agentic AI to push sequence lengths far beyond what they tested: “In the world of agentic AI, the average sequence length will grow and grow, so I’m expecting this problem to get worse.”
The Market Has Already Voted
The commercial signals are unambiguous. Intel has sold out of server CPUs through at least the end of the year. AMD has doubled its server CPU forecast. Arm and Qualcomm have both announced new CPUs designed to accelerate agentic AI. Even Nvidia — the company whose GPUs defined the AI era — has prioritized Vera, its Arm-based CPU for agentic AI, as part of the Vera Rubin platform.
Kimball reads these developments as an “absolute tell” that the industry now considers CPU performance a key part of any agentic AI system design. But he also warns of what follows: broader CPU shortages and price increases, echoing the crunches already seen with GPUs and memory. Intel has reportedly cut production of client CPUs in favor of server CPUs, even as its new 18A process has grown client-segment sales — chipmakers, like everyone else, following the money.
Why This Matters
The CPU comeback is more than a procurement headache. It’s a correction to a decade of assumption that AI progress maps cleanly onto parallel accelerator throughput. Agents don’t just think — they act, and acting means scheduling, parsing, calling, waiting, and verifying. That work has to live somewhere, and it lives on cores.
For infrastructure planners, the implication is direct: capacity models built around GPU counts alone are now incomplete. For chipmakers, an enormous new demand surface has opened. And for anyone tracking where the next shortage — and the next price spike — will hit, the answer may not be the chip inside the AI accelerator, but the one sitting next to it.