Billions a Year, Classified: The True Cost of the NSA Testing Frontier AI Models
Classified estimates shared with lawmakers reveal the NSA is spending billions of taxpayer dollars a year red-teaming frontier AI models — and the bill is reshaping the debate over who should pay for AI safety.
The most important number in the American AI policy debate this week is classified. According to a report published Thursday by NOTUS and The Washington Sun, the National Security Agency told lawmakers it is spending billions of dollars in taxpayer funds this year evaluating and testing advanced artificial intelligence models — a price tag described by two sources familiar with classified intelligence estimates as significantly greater than previously known, and one that is quietly redrawing the battle lines over who should pay for AI safety.
What the reporting actually says
The story, authored by Jeff Stein and published on September 24, 2026, rests on two sources familiar with classified intelligence estimates. Its core claims are straightforward but consequential:
- The NSA’s Artificial Intelligence Security Center (AISC), created to serve as the government’s focal point for AI security, has been systematically testing frontier models to identify potential national security vulnerabilities.
- The cost of that work has reached billions of dollars per year, funded out of classified portions of the federal national security budget. The exact figure spent to date is not clear.
- Lawmakers briefed on the numbers now believe a more comprehensive AI regulatory system could cost the government tens of billions of dollars per year — an order of magnitude that has startled Capitol Hill and brought new urgency to the debate over the government’s role in reviewing models built by some of the most resourceful companies in the world.
The political context matters. President Donald Trump has mostly resisted calls for greater federal oversight of AI, rejecting regulatory proposals from both members of Congress and the frontier labs. But after a string of high-profile hacking and security incidents, the NSA’s testing program expanded into its current form — and its price tag is now doing the arguing that pro-regulation advocates could not.
Why testing frontier models costs this much
The report identifies two dominant cost drivers, and both are structural rather than incidental.
Compute is the first. One source said the biggest AI-related cost to the NSA thus far has been extra computing power — the processing on chips necessary to run and test the models at scale. Purchasing compute has become extraordinarily expensive worldwide amid surging demand and insufficient chip production. For scale, the report notes Anthropic has secured a compute deal that could cost as much as $517 billion, according to The Information. When the government wants to rigorously exercise a frontier model — probing for cyber capabilities, dangerous knowledge, failure modes — it is bidding for the same scarce GPUs and TPUs against price-insensitive buyers.
People are the second. Top AI engineers have commanded compensation packages reaching into the hundreds of millions of dollars, far beyond anything the government can match. The NSA is widely regarded as having the deepest roster of technical talent within government, and even it may not be able to compete with frontier-lab pay packages. “Having in-house AI evaluation capability I think is extremely important and necessary. But it is genuinely expensive,” said Nathan Calvin, general counsel at Encode, an AI advocacy organization. “You’re competing in bidding with some of the most price-insensitive customers.”
The contrast with official cost estimates
What makes the classified number politically explosive is how it compares with the price tags Congress has formally budgeted for AI oversight:
- The AI Security and Innovation Act, a bipartisan House proposal to establish a center on AI risks, was scored by the Congressional Budget Office at roughly $20 million per year.
- A separate House proposal for a new AI reporting and tracking system calls for $36 million in total over five years.
Set against a classified reality of billions per year at the NSA alone, those public figures now look like rounding errors. The gap between the CBO’s arithmetic and the intelligence community’s actual experience is itself a finding: the government has been pricing AI oversight as if it were a paperwork exercise, when in practice it has become a compute-and-talent program rivaling the scale of the industry it monitors.
The Pentagon declined to comment on what the NSA is spending. “For security reasons, the Department does not discuss the technical architecture or resource allocation for its AI tools,” a Defense Department spokesperson said — while noting that the department’s annual budget approaches $1 trillion, and that some already-allocated military funds could be shifted into AI-related spending.
Three models for who pays
The report lands in the middle of a fast-moving argument about institutional design, and three competing models are now visibly on the table.
Self-regulation. Some tech leaders, including Elon Musk, have suggested the frontier labs review each other’s models prior to release without government oversight. Notably, The Information reported that Google, OpenAI and Anthropic are jointly working on a plan to create their own AI safety-focused “standards body.” Some AI safety experts reject this model outright as effectively allowing the firms to police themselves — a critique that carries particular weight the same week the industry’s own court filings have drawn fresh scrutiny.
Taxpayer-funded government testing. This is the status quo, run through the NSA’s AISC on classified budgets. The new reporting suggests it is far more expensive than anyone outside the classified world assumed, and it raises an obvious fairness question: why should the public bear the cost of certifying products built by the best-capitalized companies on Earth?
An AI industry levy. The third option, gaining traction among experts quoted in the piece, is for the government to levy a tax or assessment on AI companies to fund safety operations — both to guarantee independence and to shift the financial burden off taxpayers. Nat Purser, director of U.S. Policy at the AI Verification and Evaluation Research Institute, framed the case: “If we want the government to be equipped to assess these systems, we need to fund the computing resources and expertise that requires. I think of these tests as public goods, but there are reasonable questions about if taxpayers should foot the bill. An assessment on frontier developers could be required to help fund these independent audits, including the necessary computing resources.” Notably, Anthropic and OpenAI have signaled they want greater federal oversight and may be open to paying for it.
Why this matters beyond Washington
The classified price tag reframes three debates at once.
First, it punctures the idea that AI safety testing is cheap. Rigorous evaluation of frontier models requires running them — a lot — under adversarial conditions, with expert humans in the loop. As Purser noted, “people often underestimate how much computing power it can take to rigorously test advanced AI systems.” If evaluation is a compute problem, then whoever controls compute controls oversight, and the government’s dependence on scarce chips becomes a governance vulnerability in its own right.
Second, it changes the politics of the industry’s self-regulation push. A voluntary standards body run by Google, OpenAI and Anthropic competes directly with a taxpayer-funded NSA capability that is already spending billions. The question legislators will now ask is not whether the labs’ body is well-intentioned, but whether the public should subsidize the alternative — or whether an assessment on frontier developers should fund genuinely independent audits instead.
Third, it foreshadows the budget fights to come. If a comprehensive federal AI regulatory system truly costs tens of billions per year, that money must come from somewhere: repurposed defense funds, new appropriations, or an industry levy. Each path has a different constituency and a different set of winners. The Brookings Institution estimated this week that the U.S. AI buildout will hit $10.3 trillion; a safety regime measured in single-digit billions suddenly looks less like a constraint on that buildout and more like its insurance policy.
The bottom line
A classified number has done what months of advocacy could not: made the cost of AI oversight concrete. The NSA is already spending billions a year testing frontier models, lawmakers who have seen the estimates believe a full regulatory system would cost tens of billions, and the CBO’s public scores — $20 million here, $36 million there — describe a world that no longer exists. The next fight in AI policy will not be about whether testing happens, but about who pays for it: taxpayers, shareholders, or some combination yet to be designed.