156 Million Tokens for 800 Lines: SonarSource Puts a Price on the Coding Agent 'Context Tax'
SonarSource instrumented its own coding agent and found one ordinary 800-line PR burned ~156M context tokens and ~$41 — almost all of it re-billed cache reads from grep-and-read navigation.
Every developer running a coding agent has felt it: the bill grows with the size of the codebase, not the size of the change. On September 1, 2026, SonarSource finally put hard numbers on the phenomenon, publishing detailed telemetry from its own codebase under the name “the context tax.” The headline figure is startling. A single, ordinary ~800-line pull request — a change a human could read in five minutes — consumed approximately 156 million context tokens and cost about $41 in model fees.
What SonarSource measured
The measurement comes from an unusually credible source: SonarSource instrumented its own engineering process while building SemSitter, the semantic code-navigation engine that powers its new Sonar Vortex product for AI coding agents. Because the team keeps full traces of every agent session, the company could measure the tax precisely rather than estimate it.
The anatomy of that one pull request, which taught SemSitter’s Python analyzer to do call-site resolution:
| Metric | Value |
|---|---|
| Model round-trips in the session | 512 |
| Context window at its peak | 458,700 tokens |
| Fresh input tokens | 106k |
| Cache-read tokens (the re-billed transcript) | 152.8 million |
| Cache-write tokens | 3.1 million |
| Output tokens | 289k |
| Total context tokens billed | ≈ 156 million |
| Approximate cost of the session | ≈ $41 |
The line that matters is cache-read: 152.8 million tokens. Output was only 289,000 tokens — the agent was billed more than 500 times what it actually wrote.
Why the bill explodes: context is a tax you pay every turn
The mechanism is structural, not exotic. A coding agent does not read a file once. On every new step, the model is re-sent the entire conversation so far as input. Prompt caching makes those repeated tokens cheap per unit — roughly 10% of the input price — but you still pay for them on every single turn.
So the true cost of a token is not its size. It is its size multiplied by the number of turns it survives. Read a 600-line file on turn 40 of a 512-turn session and you have not paid for 600 lines. You have paid for 600 lines × ~470 more turns.
SonarSource traced one over-read end to end. Early in the PR, the agent needed to understand a ~67-line helper function. To find it, it read the whole 618-line file — 6,472 tokens instead of the ~700 it actually used. That read entered the conversation around turn 42 and stayed for the remaining 470 turns: 5,770 wasted tokens × 470 turns ≈ 2.7 million tokens of pure waste, about $0.54 at current cache-read pricing, from one unnecessary file read.
Fifty-four cents sounds trivial. But that PR repeated the pattern roughly ten times, plus dozens of blind tree-wide greps — several of which returned nothing and forced a second, wider grep. Across 18 comparable single-ticket PRs in the same repository, the averages were worse: ~234 million context tokens per PR, ~$65 per PR (median ~$52), ~700 model round-trips, and context windows routinely peaking between 450k and 975k tokens — brushing the 1M ceiling, at which point the agent is forced to compact and loses earlier context entirely.
The second, quieter failure: grep misses
There is a correctness cost hiding underneath the token cost. In a repository too big for the context window, grep does not just cost tokens — it misses. A regex finds the strings you thought to search for, not the call that reaches your function through an interface, an alias, or another programming language.
The SonarSource post documents a perfect example from its own codebase. A method named resolve_return_type exists in several backends — Python, TypeScript, Java, Rust, C#, and a shared core. When the agent grepped for it, the regex could not say which definition the call actually binds to, so it opened files and read generous slices to guess. Worse, the equivalent function in the C# backend is named resolve_type_node — a name no grep for resolve_return_type will ever surface. Cross-language refactors, one of the most expensive things you can ask an agent to do today, become follow-the-edge lookups instead of multi-syntax grep campaigns.
Missed call sites become failed builds, another round-trip to CI, and more rework — each with its own fresh context tax.
The proposed fix: navigate a graph, not the filesystem
SonarSource’s answer is SemSitter, a navigation engine that builds a Unified Dependency Graph (UDG) of the codebase: every function, method, class, field and parameter is a node, and the relationships between them are typed edges — calls, references, returns, has-param, is-type, contains, extends. Instead of “which files mention this string?”, the agent asks the graph “give me the definition this call binds to, the type that owns it, its return type, and its callers” — and gets back exactly that: one method body plus the answering edges, with no surrounding file and nothing to widen.
The graph also carries links beyond code-to-code. Code nodes connect to the specific documentation that governs them (documented_by edges), and documentation, tickets and design notes link to each other by meaning — so an agent can follow “this rule is refined by that ADR” without a full-text search returning fifty near-misses.
Earlier Sonar benchmarks (published June 30, 2026) reported semantic code graphs reducing agent costs by up to 36% in controlled comparisons. It is worth stating plainly: these are the vendor’s numbers, measured against the product the vendor sells. Weigh them accordingly. But the measurement methodology published this week — full traces, per-turn token accounting, named prices — is more transparent than most.
Analysis: the margin moved to navigation
Strip away the product pitch and the finding stands on its own. The bottleneck for AI coding agents in real, large codebases is not reasoning — it is navigation. Navigation-by-grep carries two compounding costs that do not show up until you measure:
- Token cost. Every blind read is re-billed on every later turn. On one ordinary PR that was 156M context tokens and ~$41; across a batch it averaged ~$65 a PR.
- Correctness cost. grep finds strings, not meaning. What it misses becomes rework and extra CI trips, each paying the tax again.
There is a memorable irony in SonarSource’s own account: the agent in the 156M-token PR was literally building call-site resolution — the exact capability it lacked for navigating its own code. Lacking an index, it fell back to grep and whole-file reads, at a 500× input-to-output ratio.
The generalization extends well past code. Any automation that scans an entire shared drive to answer one question, or re-sends a growing transcript through a model on every step, pays the same tax with a different invoice. The fixes are boring and effective: narrow the retrieval, allowlist the reads, compact deliberately, and above all measure — because the cost of a token is determined by how long it survives in the loop, not by how big it is.
If your agents work in a codebase bigger than their context window, SonarSource’s message is blunt: this tax is already on your bill. The only question is whether you have looked at the line item yet.