← All posts / Policy

The Ship That Wasn't Carrying Nukes: Inside the AI Hallucination That Almost Started a US-China War

CNN reveals a US Special Operations analyst's AI chatbot hallucinated a nuclear cargo manifest for a Chinese vessel this spring — armed boarders were ready, planes were airborne, and the strike was called off only when officials traced the report back to the bot.

Military aircraft were already in the air. Armed service members were preparing to board a Chinese-flagged merchant vessel in the Middle East. The intelligence was unambiguous: the ship was transporting components of a nuclear weapons program. And then, in the final hours before the operation, officials digging into the provenance of that intelligence discovered it was “entirely false” — a fabrication assembled by an AI chatbot.

That is the account CNN published on September 18 in an exclusive by Katie Bo Lillis and Zachary Cohen, based on four sources familiar with the episode. The report, which circulated across the US military this spring in the middle of the war with Iran, had set off alarm bells precisely because it looked so definitive. It took a last-minute provenance check to stop an operation that, as one source put it, “almost started a war.”

How a hallucination became a targeting-grade intelligence product

The mechanics of the failure are as instructive as the near-miss itself.

A Special Operations Command analyst — working with reporting that originated from US Special Operations Command Pacific in Hawaii — queried a chatbot about intelligence concerning the ship’s cargo manifest. The bot fused open-source data with classified signals intelligence held by the government, and reached a fateful, wrong conclusion about what the vessel was carrying. CNN could not determine what the cargo actually was.

Then came the second, quieter failure. The analyst used AI again to package the erroneous findings into a standard-format intelligence report — the kind of polished, familiar document that military officials have been trained for decades to trust — and disseminated it across command channels. The hallucination wasn’t just produced by AI. It was laundered into institutional trust by AI.

It remains unclear whether the chatbot was a commercial product or a government-built tool. But a former senior US official familiar with the AI systems used by military and intelligence analysts offered a blunt assessment: “The internal tools are mostly just copies of the commercial stuff wearing lipstick.”

The abort

According to CNN’s sources, the response machinery moved fast. The US military drew up plans to intercept the vessel. Armed personnel prepared to board. Military planes went airborne. Any US operation against a Chinese vessel carries the inherent risk of spiraling into armed conflict between the world’s two largest economies — which is precisely what made the eventual discovery so sobering.

Only just before the planned operation did officials dig deeper into the report’s origins and find it had been generated with AI assistance, and that the underlying cargo identification was wrong. The operation was aborted. US Special Operations Command Pacific and the Pentagon did not respond to CNN’s requests for comment.

Acceleration meets verification debt

The episode lands at an awkward moment for the Pentagon. In January, Defense Secretary Pete Hegseth released the Department’s “Artificial Intelligence Acceleration Strategy,” promising to “unleash experimentation, eliminate bureaucratic barriers” and put “America’s world-leading AI models directly in the hands of our three million civilian and military personnel, at all classification levels.”

The strategy’s ambition is now colliding with an uncomfortable structural reality that CNN’s reporting exposes: the effort is decentralized. Different parts of the government use different tools, under different orders, with different safety standards. There is no single set of standards for verifying information generated by these systems, and the reliability of the constellation of AI deployments varies widely.

Sources told CNN the military is rapidly expanding AI’s role in targeting specifically — an area where errors are fatal by definition. “AI in targeting is definitely something that is ramping up and there is no real guidance for how having a human in the loop will prevent civilian casualties or fratricide,” one source familiar with current policies said.

And the Chinese ship incident was not an isolated case. One source said this kind of hallucination has not been singular across the intelligence community since these tools proliferated through government. Younger analysts — digital natives fluent in chatbots — are particularly likely to trust the outputs uncritically, while the production pressure AI enables pushes everyone toward faster dissemination. “AI allows you to get to a bad idea faster,” one source said.

The real AI safety question, arriving early

For weeks, Washington has been consumed by a different AI debate: the theoretical scenario in which a frontier model breaks free of human constraints, a worry amplified by dire warnings from Silicon Valley engineers and executives. The Chinese ship episode reframes the conversation around a far more immediate risk — not machines acting autonomously, but humans making catastrophic decisions on the basis of confident, plausible, machine-generated falsehood.

The distinction matters for policy. Existential-risk framing invites long-horizon governance: evals, safety institutes, international coordination. Verification debt is a boring, structural problem: provenance tracking on AI-assisted intelligence products, mandatory disclosure of machine involvement in analytic workflows, training that teaches analysts that a well-formatted summary is not evidence, and verification staffing that scales with generation speed.

In an odd convergence, OpenAI, Anthropic and Google DeepMind have spent this same month in talks about forming an industry safety standards body — covering independent testing, cybersecurity and incident reporting. Whatever comes of that effort, the Chinese ship incident is a reminder that the most dangerous near-term AI failures may not happen inside the labs at all. They happen at the seam where a language model’s fluency meets an institution’s trust in paperwork.

Jake Steckler, a research scholar at GovAI and US Army veteran, put it to TechCrunch in measured terms: the incident should be “a call to add more safeguards to AI, not a reason to avoid it,” but “prioritizing adoption speed over all else will likely lead to incidents that only make service members lose trust in these systems, which ultimately is only going to slow adoption.”

What to watch

Three threads worth following as this develops:

  • Disclosure norms for analytic AI. Watch whether the intelligence community moves to require machine-involvement labeling on AI-assisted products — the cheapest fix that would have flagged this report immediately.
  • The targeting guidance gap. CNN’s sources say there is “no real guidance” for human-in-the-loop targeting safeguards. Expect congressional interest in codifying some.
  • Commercial model governance. If the chatbot turns out to be a commercial product adapted for classified use, expect fresh scrutiny of how vendors’ consumer-grade reliability claims map onto national-security workloads.

The war with Iran may end. The three-million-person AI deployment Hegseth’s strategy envisions is just beginning. The gap between generation speed and verification speed is where incidents like this one live — and until that gap closes, every well-formatted intelligence summary deserves a second look at its sources.