← All posts / Policy

'Competitors Grading Your Homework': Musk Proposes Cross-Lab 'Test Harness' Peer Review — and China Calls the Slowdown 'Fear Mongering'

At the All-In Summit, Elon Musk proposed that xAI, OpenAI, Anthropic, Google, Meta and 'three or four' leading Chinese labs run rival-built 'test harnesses' on each other's frontier models before release — a private-sector safety regime for an industry whose leaders spent the weekend begging for a slowdown that Trump now calls a 'hoax' and Beijing dismisses as 'fear mongering.'

'Competitors Grading Your Homework': Musk Proposes Cross-Lab 'Test Harness' Peer Review — and China Calls the Slowdown 'Fear Mongering'

Three days into the AI industry’s most unified safety push in years, the question is no longer whether frontier labs should slow down — it’s who, exactly, is allowed to check their work. Elon Musk offered his answer on Monday at the All-In Summit in Los Angeles: let the rivals grade the homework.

The proposal

Speaking on stage, Musk called for the top AI companies to work together and test each other’s models before they’re released to the public. The list he named is effectively the entire frontier: his own xAI — which SpaceX acquired in February 2026 — plus OpenAI, Anthropic, Google, Meta, and “three or four of the leading Chinese companies.”

The mechanism is what Musk calls a “test harness”: each lab would let its competitors run standardized evaluation suites against pre-release frontier models to probe for dangerous capabilities and safety failures.

“Instead of grading your own homework, you would at least have competitors grading your homework and raising the alarm if they see concerns.”

Musk was candid about the limits of the idea. Peer review “may not be a perfect solution,” he conceded — but “the odds that you will find issues are dramatically greater.” He also acknowledged that rival labs haven’t agreed to the proposal, which for now remains exactly that: a proposal. “What I’m suggesting here is it’s a step in the right direction and it’s something that we do quickly,” he said, adding, “I think it’s probably something that China would agree to.”

The framing matters. Rather than a government-mandated licensing regime or an international treaty, Musk is pitching a horizontal accountability structure: labs policing labs, with commercial rivals — the parties with the strongest incentive to find each other’s weaknesses — acting as the first line of defense.

Why now: a weekend that upended the debate

Musk’s proposal didn’t land in a vacuum. It came at the end of a stretch in which the AI industry’s own leaders made their most coordinated appeal yet for restraint:

  • Anthropic CEO Dario Amodei published an essay proposing to “slow the pace” of frontier AI development, suggesting labs could accept third-party assessments from “embedded evaluators” who would verify safety practices from the inside.
  • Musk and OpenAI CEO Sam Altman both publicly backed Amodei’s slowdown call — a rare show of unity between two CEOs with a long, bitter history.
  • The spark for much of it was Jacob Coxon, a researcher who quit Anthropic (after previously working at OpenAI) and wrote that the leading labs are “gambling with our lives.” Anthropic alignment lead Evan Hubinger amplified the warning: “Jacob is correct here — we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.”

Against that backdrop, Musk’s test-harness idea reads as an attempt to translate vague slowdown rhetoric into something concrete and fast. Amodei’s “embedded evaluators” need to be recruited, accredited, and trusted; a rival-run test harness needs only a phone call between competitors — at least in theory.

Washington’s answer: ‘hoax’

The White House, however, wants no part of it. President Trump spent Monday posting a flurry of messages on Truth Social calling fears over AI a “hoax” and a “scam,” explicitly rejecting calls for greater regulation. He argued that any attempt to pace AI development would hand China the upper hand — and, according to one account, quipped that “the only guardrail AI needs is a high IQ president.”

National Economic Council Director Kevin Hassett told CNBC on Tuesday that the private sector is the “right place” to address AI concerns, saying the government will keep monitoring the industry and will use “law enforcement when necessary to make sure that the firms are acting responsibly.”

That position contains a deep irony: an administration refusing to regulate AI is effectively endorsing — by omission — exactly the kind of industry self-policing Musk is proposing. If Washington won’t build guardrails, the guardrails will be built by competitors probing each other’s models, or not at all.

Speaker Mike Johnson echoed the “hoax” line at a Tuesday news conference, telling reporters the warnings were “even wilder than the RUSSIA, RUSSIA, RUSSIA HOAX, or the Global Warming Scam.”

Beijing’s answer: ‘fear mongering’

China went further. A spokesperson for China’s Foreign Ministry on Monday called the AI companies’ push for a slowdown “fear mongering,” according to a Reuters translation of the remarks.

That response all but torpedoes the international dimension of Musk’s plan. His proposal presumes that “three or four of the leading Chinese companies” would submit their frontier models to inspection by OpenAI, Anthropic, and the U.S. intelligence-adjacent establishment — and that Washington would tolerate Chinese labs running test harnesses on American frontier models in return. Musk’s own guess (“I think it’s probably something that China would agree to”) now sits against Beijing’s official characterization of the entire slowdown movement as scare tactics.

It also sharpens the dilemma Amodei himself acknowledged over the weekend, calling China the “toughest dilemma” for his proposed slowdown: unilateral restraint by U.S. labs is unverifiable and unreciprocated, while mutual verification requires a level of U.S.–China trust that currently does not exist. Mid-September AI safety talks between the two countries — including proposals for labs to jointly “police” AI-directed cyberattacks — remain tentatively scheduled and unconfirmed.

Analysis: adversarial auditing, without the auditor

Strip away the geopolitics and Musk’s proposal is a genuinely interesting governance design — arguably the most pragmatic on the table.

The incentive argument is strong. Self-evaluation suffers from an obvious conflict of interest: no lab wants to find the result that delays its launch. Competitors have the opposite problem — they are maximally motivated to find each other’s flaws, and they possess the deepest technical talent to do it. Formalizing that rivalry into a pre-release testing protocol converts market competition into a safety resource.

But the trust problem is fatal in cross-border form. A test harness run by a rival requires sharing pre-release model access with an adversary — handing over capabilities to probe, reverse-engineer, and benchmark against. U.S. labs sharing frontier access with Chinese labs (or vice versa) touches export controls, national security review, and espionage law. The domestic version — xAI testing OpenAI’s models, Anthropic testing Google’s — is more plausible, and echoes the quiet frontier-standards cooperation already reported among Anthropic, OpenAI, and Google. But it would cover only half the world’s frontier.

And it’s no substitute for what the debate is actually about. The weekend’s warnings were about pacing — whether to train the next generation of models at all. A test harness evaluates whatever you’ve already built; it doesn’t slow down the building. Musk, who has simultaneously insisted that “nothing can shut down open source,” is in effect offering a way to keep racing — with better brakes inspected by the other racers.

That may be precisely why it’s the one proposal with a chance: it asks no one to stop, requires no treaty, and could start with a handful of phone calls. Whether any lab actually picks up the phone — in a week when the U.S. president calls the risk a hoax and China calls the worry fear mongering — is the real test.