← All posts / Industry

Glean Claims Claude Cowork Costs 5x More Per Task: Inside the Benchmark Wars Reshaping Enterprise AI

Glean's benchmark puts its auto-routed assistant at $0.58 per enterprise task versus Claude Cowork's $2.98 — an 81% gap it attributes to context architecture, not model quality. The Information amplified the claim this week, and it lands amid Uber and ServiceNow blowing their annual AI budgets in months.

Glean Claims Claude Cowork Costs 5x More Per Task: Inside the Benchmark Wars Reshaping Enterprise AI

The sharpest attack on Anthropic’s enterprise ambitions this week didn’t come from a rival model lab. It came from a partner. Glean — the enterprise search and work-AI company whose entire product sits on top of other people’s models — published a benchmark claiming that its assistant completes typical enterprise tasks for $0.58 apiece, while Claude Cowork averages $2.98 on the same work. That is an 81% gap in token costs, and The Information’s Kevin McLaughlin distilled it into a headline guaranteed to reach every CFO with an Anthropic contract: “Anthropic Customers’ Bills Are 80% Higher Than They Need to Be, Glean Says.”

The claim lands at a moment of genuine anxiety about AI spend. Uber disclosed that it burned through its entire 2026 AI coding budget in four months. ServiceNow reported the same fate weeks later, with its CIO calling it “a really hard problem.” Glean has a name for the era that preceded this reckoning — “tokenmaxxing,” the practice of treating token consumption as a productivity metric — and its blog post reads as both a eulogy for that era and a sales pitch for what replaces it: routing, context, and right-sized intelligence.

What the benchmark actually measured

The setup is straightforward. Glean ran its Assistant with auto routing enabled against Claude Cowork running Claude Sonnet 5 with high reasoning — Anthropic’s own recommended configuration for everyday work — across more than 180 enterprise tasks spanning sales, engineering, marketing, HR, product, finance, and customer support. Claude Cowork was given a fair harness: off-the-shelf MCP connectors for Google Drive, Gmail, Google Calendar, Slack, Atlassian Rovo, Linear, Intercom, and Sigma, plus local MCP servers for Salesforce, GitHub, and GCP.

The headline numbers:

  • Cost per task: Glean $0.58 vs. Cowork $2.98 — an 81% reduction
  • Total tokens: Glean 1.3 million vs. Cowork 4.4 million — a 70% reduction
  • Human preference: graders preferred Glean’s answers 78% of the time on overall preference, correctness, completeness, and interaction quality

Notably, this is the second time Glean has run this comparison. A mid-August version of the benchmark, which made the rounds on LinkedIn, showed $0.45 versus $1.84 per task — a 4x gap. The new numbers are worse for both sides in absolute terms but roughly consistent in ratio, which suggests the comparison isn’t a one-off cherry-pick but a repeatable measurement of architectural difference.

Why Glean says it wins: context, not models

The most interesting part of the benchmark is the explanation, because it has nothing to do with having a better model. Glean doesn’t have a frontier model. Its argument is that the cost difference comes from architecture:

Pre-indexed enterprise context. Glean starts from a unified, permission-aware, ranked index of company information. Claude Cowork’s federated MCP approach has to search each connected system individually — often overfetching results, then burning tokens to normalize data, resolve conflicts, and repeat reasoning loops it shouldn’t need. If the answer to “what did the EMEA team decide about pricing last quarter” requires five tool calls through Slack, Drive, and Salesforce instead of one retrieval, the token meter runs the whole time.

Harness design. Glean keeps tool outputs and intermediate state in sandbox files rather than reloading everything into the context window on every turn. It progressively loads only the tools and schemas a task needs, and isolated sub-agents keep unrelated context out of the main reasoning loop. This is, quietly, an indictment of how most agent harnesses are built: the default pattern of stuffing every intermediate result back into context is the single largest driver of per-task cost.

The quality result — 78% preference — is arguably more damaging than the cost result, because it removes the obvious rebuttal. If Cowork were cheaper but worse, Anthropic could argue you get what you pay for. Glean’s claim is that better context retrieval produces both cheaper and more preferred answers.

The Pareto frontier: no single model wins

Alongside the head-to-head, Glean published a Pareto frontier analysis across 1,000 enterprise tasks, 37 model-and-reasoning configurations, 11 model families, and 81 head-to-head matchups, scored with Bradley–Terry methodology into a “Glean quality score.” The findings are worth sitting with:

  • GPT-5.6 Luna (xhigh) sits on the frontier: quality score 55 at $0.0819 per task
  • GLM 5.2 (high) — open source — scores 57 at $0.3489
  • Gemini 3.7 Flash (high) scores 61 at $0.4748
  • Kimi K3 (high) scores 63 at $0.8995
  • Claude Opus 5 (high) anchors the expensive end: 67 at $2.9605

The spread: the cost gap between the cheapest and most expensive frontier models is 36x, for a 22% quality-score gain. And models like Kimi K3 and Gemini 3.7 Flash capture most of Opus 5’s quality advantage at 3–6x lower cost. Glean’s conclusion — “there’s no single best model” — is also its product thesis: its auto routing uses a small specialist model called Glean Waldo, post-trained on NVIDIA Nemotron 3 Nano, to decide at runtime which of 40+ models and which reasoning effort a task deserves.

The timing is not accidental. The past month brought new open models — GLM 5.2, Kimi K3, DeepSeek V4 Flash — at aggressive price points, plus OpenAI’s July cuts to GPT-5.6 Luna and Terra. Enterprises responded by mixing open models into their stacks, and inference platforms serving them are gaining share. A benchmark that shows open models on the cost-quality frontier is a benchmark that sells routing.

Botsitting: the unbudgeted tax

The post’s most quotable statistic isn’t about tokens at all. Glean’s Work AI Institute surveyed 6,000 full-time digital workers and found they spend 6.4 hours a week “botsitting” — feeding AI context, supervising output, and cleaning up mistakes. That is more time than they spend using AI to produce actual work. Grader feedback from the benchmark traced this correction labor to concrete failure patterns: misread tasks, outputs hard to read or share, missing role or account context, incomplete or stale evidence, and unsupported conclusions.

The implication is uncomfortable for every vendor in the space: if the model needs six hours of weekly supervision per employee to be usable, the labor cost dwarfs the token cost. Glean’s argument is that starting with the right context already in place eliminates most of that correction work.

Read the fine print

Healthy skepticism applies. This is a vendor benchmark, run by Glean, on synthetic queries against Glean’s own production data, with Glean picking the comparison configuration. Claude Cowork on Sonnet 5 is a single point in Anthropic’s configuration space, and Anthropic would surely argue that Opus-class models or different harness setups close the preference gap. The 78% preference figure comes from Glean’s own graders. And Glean has an obvious commercial interest: it launched Glean Tau, a desktop AI workspace, the same week as the benchmark, and every routing decision its platform makes is a line item it monetizes.

But even discounted for self-interest, the core finding matches what enterprises are discovering on their own invoices. The economics of agentic AI are dominated not by which frontier model you pick, but by how many tokens your architecture burns retrieving context the system should already have. Anthropic’s response so far has been structural rather than rhetorical — aggressive cache-read pricing cuts in Fable 5.1, usage-based enterprise billing — which suggests it reads the same writing on the wall.

The benchmark wars of 2026 are no longer about who has the smartest model. They are about who can prove, with numbers, that their route to the same answer costs less. Glean just fired one of the loudest shots.