Google DeepMind Runs the World's First Double-Blind AI Evaluation: Neither Side Could Peek
DeepMind, AVERI, OpenMined and MLCommons evaluated Gemini 2.5 Flash-Lite inside a cryptographic enclave — the evaluator couldn't see the weights, Google couldn't see the test prompts, and benchmark contamination became physically impossible.
Imagine a student about to take a high-stakes exam. If they accidentally glimpse the test questions in advance, a perfect score stops meaning anything. That is precisely the dilemma the AI industry faces every time a frontier model is sent to external evaluators — and on August 27, 2026, Google DeepMind announced that it has finally solved it. Together with the AI Verification and Evaluation Research Institute (AVERI), OpenMined, MLCommons, and the Singapore AI Safety Institute, DeepMind has completed what it calls the world’s first double-blind evaluation of a proprietary, frontier-class AI model: an audit in which the evaluators could not see the model, and the model’s owner could not see the test.
The contamination problem
The announcement, authored by DeepMind’s William Isaac, Sol Messing and Kristian Lum, addresses one of the most quietly corrosive issues in AI measurement: benchmark contamination. If a model has already seen the test questions during training — deliberately or inadvertently — its scores are inflated in ways no one can easily audit. Policymakers, researchers, and enterprises all need benchmarks to reflect a model’s true capabilities and safety. When models can “peek” at the evaluation in advance, that trust erodes.
Historically, high-stakes external evaluation forced an uncomfortable trade-off. Either the evaluators handed over their testing prompts — risking the model provider absorbing them and optimizing against future tests — or the provider handed over model weights, exposing some of the most valuable intellectual property on Earth to theft and leakage. Zero-logging policies and contractual NDAs have long papered over this gap, but a promise is not a proof.
How the double-blind mechanism works
The pilot, conducted in July and August 2026, eliminated the compromise with hardware. The evaluation ran inside a secure enclave — a trusted execution environment (TEE) configured using Confidential Space within Google Cloud’s Confidential Computing portfolio. The hardware cryptographically attests to exactly which code will run before either party’s assets enter the environment; both sides review and approve that code. Then the enclave executes it, releasing only the agreed outputs to the agreed recipients.
The concrete arrangement, as AVERI’s pilot report describes it, was a four-party choreography:
- Google DeepMind provided a production model — Gemini 2.5 Flash-Lite — whose weights were never exposed to any outside party.
- MLCommons supplied a never-before-used set of prompts from its AILuminate safety benchmark family, covering hazard domains including chemical, biological, radiological, nuclear and explosive (CBRNE) risks, cyberattacks, hate speech, self-harm, and violent-crime elicitation. The prompts were produced under strict non-disclosure and origination requirements and never shared broadly, even with working-group members.
- AVERI encrypted the prompts with a private key shared with no other party, jointly ran the evaluation with DeepMind using software built and adapted by OpenMined, then decrypted and graded the outputs against AILuminate criteria.
- OpenMined supplied the enclave-management software that defines which computations the enclave is permitted to carry out — a practice its authors call structured transparency: a technology-enforced middle ground between “share everything” and “share nothing.”
The result is double-blindness enforced by silicon and cryptography rather than contract language. DeepMind was technically unable to collect, store, or train against the MLCommons prompts. AVERI, OpenMined, and MLCommons were equally unable to inspect the model’s weights. AVERI delivered its findings to DeepMind as a confidential report summarizing observed success modes and failure modes with quantitative benchmark results — while the prompts and raw outputs stayed private.
From proof-of-concept to production
This was not the first time secure enclaves were proposed for AI evaluation — and the continuity matters. In a 2024 pilot, OpenMined collaborated with Anthropic and the UK AI Security Institute to demonstrate that two organizations could jointly approve and run an evaluation inside an NVIDIA H100 GPU enclave. That earlier experiment, however, used stand-ins: GPT-2 as a proxy for a frontier model, and a five-row evaluation as a proxy for a confidential biosecurity benchmark. The computation ran end-to-end in just over a minute, proving the mechanism, but nobody had risked a real asset.
The 2026 pilot crossed that line. A proprietary, production-grade Gemini model was evaluated against a genuinely confidential, genuinely unused benchmark, with privacy guarantees for both sides. It is the first demonstration that leading AI companies’ closed models can be independently benchmarked without either party surrendering its crown jewels.
Why regulators should care
The timing is not accidental. Third-party evaluation is rapidly becoming a legal requirement, not a courtesy. Under the EU AI Act’s General-Purpose AI Code of Practice, signatories must ensure qualified independent evaluators can conduct model evaluations with adequate access, information, and time. In the United States, Illinois’ SB 315 will require annual third-party audits of large frontier developers — including audits of whether they follow their own safety policies — starting in 2028, and other state and federal proposals could compel deeper access sooner. Yet these laws demand that audits protect security and privacy while saying almost nothing about how to operationalize both at once.
AVERI’s report makes the policy argument explicit: policymakers should feel encouraged by these results “to be ambitious in requiring that deep, secure access be provided,” because a policy-based demand signal will accelerate the development of enclave-based auditing technologies. Secure enclaves also open a door that has long been considered jammed shut: international verification. Scenarios like verifying compliance with an international AI agreement require exactly this kind of mutual, cryptographically enforced confidentiality between parties that do not trust each other.
The honest caveats
Both organizations are careful about what the pilot does not prove. The security design addressed many, but not all, theoretical avenues by which a developer could tamper with evaluations — higher-stakes, lower-trust settings would need additional assurances. Model evaluation is also not the entirety of AI auditing: double-blind guarantees do not yet extend to non-model components of an AI system, or to a company’s data practices, internal documents, and computing hardware. And “low-tech” auditing — humans reading documents, asking questions, and exercising judgment — will continue to play a central role. Translating enclave-based results into formal auditing standards is explicitly left as future work.
What it means
For the benchmarking ecosystem, the pilot offers a way out of an arms race in which public benchmarks decay as they are absorbed into training data. For auditors and AI Safety Institutes, it offers deeper access with fewer trade-offs. For DeepMind, it is a reputational bet that verifiable evaluation — not just favorable scores — will become the currency of trust in frontier AI. If enclave-based double-blind evaluation becomes routine, “we tested it and here’s the cryptographic proof that nobody cheated” becomes an auditable claim rather than a press release. In a year when more than 100 companies signed an open letter warning of a “limited window” to defend against AI threats, infrastructure for verified claims about model safety may be exactly what the window demanded.
Sources
- [1] https://deepmind.google/blog/piloting-the-worlds-first-double-blind-ai-evaluations/
- [2] https://www.averi.org/ourwork/averi-pilot-report-the-worlds-first-double-blind-eval
- [3] https://www.techrepublic.com/article/news-google-deepmind-gemini-tests-apac-singapore/
- [4] https://explainx.ai/blog/google-deepmind-double-blind-ai-evaluation-benchmark-contamination-august-2026