← All posts / Industry

The Cheapest Token Isn't the Cheapest Answer: AlphaSense Flips the AI Cost Story

A new AlphaSense study of 246 financial-analysis tasks finds GPT-5.6 Sol and Claude Opus 4.8 beat Chinese rivals on both quality AND total cost — because smarter models spend fewer tokens per task. The per-token price war may be measuring the wrong thing.

The Cheapest Token Isn't the Cheapest Answer: AlphaSense Flips the AI Cost Story

For months, the story of enterprise AI economics has been written in one currency: the price of a million tokens. Chinese models charge a fraction of their American counterparts — Moonshot’s Kimi K3 lists at $15 per million output tokens against $25 for Anthropic’s Opus 4.8 and $30 for OpenAI’s GPT-5.6 Sol — and buyers have responded by switching in record numbers. OpenRouter data this summer showed Chinese models overtaking Claude and ChatGPT in developer token volume for the first time.

A new study from AlphaSense, the AI-powered market intelligence platform, argues that everyone buying on sticker price is measuring the wrong thing. After running 246 real financial-analysis tasks — analyzing earnings transcripts, SEC filings, analyst estimates, and acquisition activity — AlphaSense found that the “expensive” American models beat their cheaper Chinese rivals on both quality and total cost.

The findings

The numbers are striking. On a median basis across the task set:

  • OpenAI’s GPT-5.6 Sol delivered answers roughly 20% higher in quality than Moonshot’s Kimi K3 — while costing about 13% less overall.
  • Anthropic’s Opus 4.8 scored around 13% higher in quality than Kimi K3 at roughly half the total cost.
  • Tested Chinese rivals included Kimi K3 and GLM-5.2; both lost on the cost-per-completed-task metric despite their lower per-token prices.

How can a model charging $30 per million tokens cost less than one charging $15? The mechanism is token efficiency. More capable models often require fewer tokens and fewer processing steps — fewer retries, less scaffolding, shorter reasoning chains — to finish the same job. Total spend is price × volume, and in complex knowledge work, the volume term dominates.

“Some of the more expensive models, the ones that look more expensive based on just their price per token, actually ended up being less costly because they were more efficient in using tokens,” AlphaSense CEO Jack Kokko said.

Why this matters now

The study lands in the middle of a full-blown price war. According to the Financial Times, leading US model prices have fallen nearly 25% since mid-July — the sharpest one-month drop since ChatGPT launched — as OpenAI and Anthropic fight cheap Chinese alternatives. OpenAI cut GPT-5.6 Luna by 80%; Anthropic launched Claude Opus 5 at half the price of its frontier Fable 5.

The two stories don’t contradict each other — they complete each other. Falling list prices are the supply-side response to competition. AlphaSense’s data is the demand-side correction: what you pay per token matters less than what you pay per completed task at an acceptable quality bar. An enterprise that migrated to a cheaper model and saw its agent loops triple in length, retry rates climb, and output quality slip has been paying for the discount in ways the invoice’s unit price never showed.

The routing middle ground

Notably, AlphaSense doesn’t conclude that frontier models win everywhere. The report acknowledges two durable cases for open-weight models: companies running workloads on their own infrastructure avoid per-token charges entirely, and undemanding tasks like email summarization don’t need frontier intelligence at all.

The company’s own platform models the pragmatic answer: a routing system that assigns different parts of a query to different models — a capable model to plan the response, a smaller cheaper one to execute it. The cost-optimal architecture for most enterprises will likely be a portfolio, not a pick.

What to watch

AlphaSense’s study covers one domain — financial research — and one task set of 246 items. That’s a real workload, not a benchmark suite, which cuts both ways: it’s more realistic than MMLU-style scoring, but the efficiency advantage of frontier models may be narrower or wider in other domains like coding, customer support, or document review.

Still, the direction is clear. As AI procurement matures, the metric that decides vendor selection is shifting from price-per-token to cost-per-completed-task — and on that metric, the race is far more open than the price tables suggest. US labs betting that efficiency can offset a 2x list-price premium just got their first strong piece of evidence. Chinese labs, meanwhile, now know exactly where the next front in the cost war opens: tokens needed per task, not dollars per token.

For buyers, the practical takeaway is simple: before the next renewal, benchmark on your own tasks, measure total spend and output quality together, and treat the pricing page as the start of the negotiation — not the whole story.