← All posts / Industry

Nearly a Billion Dollars for Care That Never Happened: Blue Cross Pins $942M on AI Coding Tools

A Blue Cross Blue Shield Association study finds AI coding tools and ambient scribes drove $942 million in added insurer costs in 2024–2025, as diagnosis codes surged without matching treatments — the clearest quantification yet of AI-driven upcoding.

Nearly a Billion Dollars for Care That Never Happened: Blue Cross Pins $942M on AI Coding Tools

The number landed quietly on a Thursday: $942 million. That is how much extra spending Blue Cross Blue Shield insurers absorbed over two years as hospitals and clinics adopted AI tools that find more things to bill for — without, according to the insurers’ data, actually delivering more care.

The study, released September 24 by the Blue Cross Blue Shield Association (BCBSA) and reported by Reuters, is the sharpest quantification yet of a phenomenon payers have warned about since ambient AI scribes began appearing in exam rooms: AI-assisted upcoding. And it lands at a moment when AI’s most defensible clinical use case — lifting the documentation burden off burnt-out physicians — collides with the billing incentives of the American fee-for-service system.

What the study found

BCBSA analyzed inpatient billing across its network of 31 independent licensees, which together cover more than 100 million Americans. Comparing 2024–2025 claims against a 2023 baseline, researchers tracked how often providers billed for secondary conditions — coexisting illnesses documented alongside the primary reason for admission. In hospital settings, such conditions can push a case into a higher complexity tier, which commands higher payments.

The headline figures:

  • $653 million of the increase came specifically from more frequent billing of secondary conditions during 2024–2025.
  • $942 million is the total added cost from more intensively coded care over the two-year period, versus 2023.
  • For patients undergoing major bowel surgeries, documented secondary conditions jumped sharply: partial intestinal blockages rose 55%, and acid overload in the body 33%, between Q1 2023 and Q4 2025.

The mechanism is not exotic. Providers used AI to scan existing patient records for billable secondary diagnoses, or deployed ambient scribes — tools that passively listen to the clinician–patient conversation and draft medical notes. More documented complexity means more reimbursement. A note that once read “routine admission” now arrives padded with every codable condition the model could infer from a decade of chart history.

Diagnoses without treatment

The most damning part of the analysis is what did not change: treatment rates.

“If patients are truly sicker, we’d expect to see more treatment,” said Luke Chalker, senior vice president of product and data science at BCBSA. The data showed the opposite. Dr. Razia Hashmi, BCBSA’s vice president of clinical affairs, noted that documented complexity in bowel surgery cases did not correspond with higher rates of actual treatment. Anemia diagnoses climbed, for example, but blood transfusions — the standard response to a genuine red-blood-cell deficit — did not.

“The disconnect between diagnoses and treatment suggests that AI is identifying more billable conditions, not sicker patients,” Chalker concluded.

That single sentence is the study’s thesis, and it is the pattern auditors look for when distinguishing legitimate documentation improvement from upcoding: codes without consequences, diagnoses without interventions. The AI did not make patients sicker; it made their charts more expensive.

A fight that was already brewing

BCBSA is not a disinterested party — it represents the entities that pay these bills, and payers have every incentive to litigate coding intensity in public. But the association’s September report is the second act of a story it opened in March 2026, when earlier BCBSA/Blue Health Intelligence research first tied AI scribes to rising hospital billing and estimated excess inpatient spending from AI-enabled coding in the hundreds of millions. Other insurers, including Centene, have separately said health systems’ AI tooling has driven aggressive or inappropriate reimbursement demands.

Hospitals tell a different story: that documentation has been under-coded for decades, that physi- cians under time pressure historically missed billable secondary diagnoses, and that AI simply corrects an information gap. There is truth on both margins — some of the rise likely reflects legitimate capture of real conditions. But the “sicker on paper, untreated in practice” signature is hard to explain away entirely.

Meanwhile, the policy literature has been tracking this arms race for over a year. A 2025 policy brief in PMC described ambient scribes’ billing effects as a looming “coding arms race,” and 2026 payer-side research shows insurers have begun downcoding AI-scribed claims in response — a countermeasure that risks penalizing honest documentation alongside inflated claims.

Why it matters beyond the bill

Three larger implications are worth sitting with.

First, this is agentic AI’s first fully quantified externality. The industry spends its safety discourse on frontier models and loss-of-control scenarios. Here is a mundane, measurable harm pattern: inference systems optimizing against reimbursement rules at scale, moving nearly a billion dollars across two years in one insurer network alone. Extrapolated across the US payer landscape, the true figure is a multiple of the BCBS numbers.

Second, the incentive gradient points the wrong way. Ambient scribes genuinely reduce physician burnout — that benefit is real and well-documented. But the ROI case that sold hospitals on these tools was never just about clinician wellness; it was about revenue integrity, which in practice means capturing every codable condition. When the tool’s pitch is “find more billable codes,” the hospital buys, the payer pays, and the premium-payer — the patient, ultimately — absorbs the difference in rising premiums.

Third, verification is now the bottleneck, not generation. The core asymmetry exposed here is that AI can generate a plausible secondary diagnosis in milliseconds, while verifying whether that diagnosis reflects real sickness requires human clinical judgment that the system has no room to fund. Every efficiency gained on the documentation side becomes a cost on the audit side. Expect payers to answer with their own models, and the billing-audit arms race to become AI-versus-AI.

What comes next

The BCBSA report gives regulators and insurers their cleanest evidentiary baseline yet. Likely follow-ons: payer audits specifically targeting facilities with rapid post-scribe-adoption coding shifts; state insurance commissioners folding coding-intensity metrics into rate reviews; and possibly legislative attention, given that healthcare billing is one AI application area where harm is concrete, monetized, and traceable to identifiable actors.

For the AI industry, the uncomfortable lesson is that deployment context is destiny. The same ambient-listening technology that frees a physician from typing can quietly reprice a hospital stay — not because the model is misaligned in any lab-defined sense, but because it is dutifully optimizing the objective its buyer actually paid for. The $942 million is not a failure of intelligence. It is a failure of incentives, priced per claim.