← All posts / Industry

AI Became AI's Biggest Customer: Agent Token Use Up 14x on OpenRouter Since February

OpenRouter data suggests February 6, 2026 was the last day humans out-consumed AI agents. Agent token usage has since grown 14x to 7.3 trillion tokens weekly — nearly 5x human levels — and cache economics are quietly rewriting the bill.

AI Became AI's Biggest Customer: Agent Token Use Up 14x on OpenRouter Since February

The Day the Machines Took Over the Queue

Somewhere in the logs of one of the internet’s busiest AI routing layers, a quiet milestone passed this winter. According to OpenRouter analyst Peter Walker, February 6, 2026, may have been the last day human users consumed more tokens than AI agents. Nobody celebrated it. There was no launch event, no blog post. But six months of data published in late August confirm what that crossover started: on OpenRouter, the machines are now the customers, and they are eating like it.

The numbers, first reported by The Decoder’s Matthias Bastian on August 23 and drawn from a16z’s “Charts of the Week” newsletter dated August 21, are stark. Since February 2026, weekly token consumption by AI agents on OpenRouter has jumped from 0.51 trillion to 7.3 trillion tokens — a roughly 14x increase. Human token usage over the same period grew just 2.8x. Agents now burn nearly five times as many tokens as human users on the platform.

AI, in other words, has become AI’s biggest customer.

Why Agents Eat So Much

The composition of that consumption matters more than the raw volume. Agents don’t use tokens the way people do. A human works in prompt-and-response exchanges: type a question, read an answer, repeat. An agent is built to iterate toward a goal. Its opening prompt carries a token-intensive pre-fill — policies and procedures, code guidelines, project context — and from there it reads and writes incrementally, spinning up additional AI processes along the way, re-reading its own working context as it goes.

Walker’s own summary of the shift is blunt: most tokens spent in AI work today are simply re-reading. As agents became the dominant token users, the flow of usage inverted — from generating novel output toward repeatedly ingesting accumulated context.

That structure shows up directly in the billing data. The Decoder reports that nearly 70 percent of agent token usage on OpenRouter comes from cached prompts; a16z’s newsletter, citing the same OpenRouter data, puts cached tokens at more than 85 percent of agentic token burn and credits them with nearly all of the relative growth. The two figures measure slightly different things, but the direction is identical: agentic AI is, overwhelmingly, a cache-read workload.

This has three consequences.

First, the bill isn’t rising as fast as the tokens. Cached tokens are billed at a fraction of the pre-fill rate, so the unit economics of running agents are far better than a 14x usage chart implies. That is a big part of why agent deployments haven’t simply priced themselves out of existence.

Second, memory is the new bottleneck. Caching is, by definition, memory-consumptive — which a16z identifies as a key reason high-bandwidth memory (HBM) remains in such demand across the industry.

Third, efficiency pressure is reshaping architecture. When most spend is cache reads, engineering effort moves to collapsing redundant context. PPC Land documents one media-buying setup that folded twelve separate Model Context Protocol calls into a single buyer agent specifically to contain token consumption.

One Important Caveat

OpenRouter skews toward open-weight models, which tend to be less token-efficient than the frontier models from OpenAI and Anthropic — models whose direct API traffic doesn’t flow through the router. So the absolute mix on OpenRouter isn’t a perfect mirror of the industry. But as The Decoder notes, the trend almost certainly looks similar at the major labs: token inflation began with reasoning models that think longer before answering, and agents compound it by design.

The Collateral Damage Is Already Visible

The a16z newsletter pairs the OpenRouter data with a second-order signal: traffic to the legacy automation platforms Zapier, n8n, and Make has fallen by double digits on a trailing twelve-week basis according to Similarweb. All three predate large language models and dominated workflow automation before them. The only platform in the measured set gaining traction is Gumloop, a 2023-era, agent-native builder. Site traffic is an imperfect proxy — a mature product can shed visits without shedding customers — but a synchronized, sustained decline across three incumbents is hard to dismiss. The pre-agentic automation stack is starting to look like collateral damage of the agentic one.

Meanwhile, OpenAI-attributed data in the same newsletter shows output tokens roughly doubling at the typical enterprise since April 2025 — but the top decile is up more than 17x, opening an eightfold gap with the median firm across industries, and nearly twelvefold within the tech sector itself. Adoption of agent-style tooling (plug-ins, Skills) among top-decile firms runs at two to six times typical levels. The AI economy isn’t just growing; it’s concentrating.

Even the professional categories are surprising. The fastest-moving cohort on Codex adoption isn’t software — it’s legal, up 108x since February 2026, though Sternstein flags the caveat that part of this reflects the broader Codex rollout.

The Sober Counterweights

None of this is unambiguously bullish. KPMG research published in August found 49% of surveyed leaders had scaled back agent rollouts after running costs outran delivered value, and only 35% claimed full visibility into what their AI systems cost to operate. LayerX found just 18.24% of enterprise employees use AI tools weekly, with 47% of enterprise AI conversations happening outside corporate governance. And France’s competition authority, citing Sensor Tower data, named OpenAI, Google, and Anthropic as holders of more than 84% of the global AI agent market as of May 2026 — a concentration it flagged as a lock-in risk in a July opinion.

The picture, then, is not “agents everywhere.” It is “agents pulling away where they work, at a cost structure built on cache.”

What to Watch

The OpenRouter dataset — the same routing layer Stripe agreed to acquire earlier this month, a deal that would place agent traffic metering inside one of the world’s largest payments networks — will be the place to watch these curves. Three signals matter next: whether the 14x agent growth rate holds, flattens, or accelerates as multi-agent architectures compound; whether cache pricing spreads standardize across providers as agentic traffic dominates; and whether human-facing token usage stagnates into a rounding error as interfaces themselves become agents.

The crossover that passed quietly on February 6 wasn’t a model release or a benchmark record. It was something more structural: the moment AI workloads stopped being an extension of human typing and became a machine economy with its own consumption habits. Seven point three trillion tokens a week, and climbing — mostly spent talking to itself.