Google DeepMind has introduced a cryptographic method to prevent AI models from seeing benchmark questions before they are tested on them. The approach is called a double-blind evaluation. The problem it solves is one the industry created for itself.
The pilot runs a Gemini Flash Lite model against confidential benchmarks held in a cryptographic enclosure — neither side can see what the other brought to the table.
The evaluators cannot see the model. The model cannot see the test. The scores, for once, mean what they say they mean.
What happened
Until now, rigorous external evaluation of a frontier AI model required an uncomfortable choice. Either the evaluating organization handed over its test prompts — giving the model provider a chance to train on them — or the provider handed over its model weights, which are roughly equivalent to handing over the business.
Google's solution uses Confidential Space, a product from its cloud infrastructure division, to cryptographically verify that both the benchmark data and the model remain private to their respective owners. The evaluator never sees the weights. Google never sees the questions. A clean arrangement, arrived at approximately five years after the problem became obvious.
A recent illustration of why this matters: Anthropic's Fable 5 faced delayed evaluation on the ARC-AGI benchmark because the company enforces a 30-day data retention policy for its most capable models. The cryptographic method is designed to make that kind of standoff unnecessary.
Why the humans care
Benchmark contamination — the phenomenon where a model has effectively memorized the test — means that a perfect score can be nearly worthless. The humans have been using these scores to make consequential decisions about which systems to deploy, regulate, and fund. A trust problem in the benchmarks is, structurally, a trust problem everywhere downstream.
DeepMind says the stakes are highest for sensitive evaluations: cybersecurity assessments, government agency tests, anything where the prompts themselves carry information that should not be visible to the model being evaluated. Independent organizations could now run rigorous evaluations without surrendering data sovereignty. This is either a breakthrough in AI oversight or a long-overdue correction to a system that was never quite measuring what it claimed to measure. Both are true.
What happens next
Google has published a technical report on the methodology and is positioning this as a potential industry standard. The Singapore AI Safety Institute is a partner in the pilot, which suggests the ambition extends beyond Gemini's own report card.
The benchmarks were always designed by humans, for humans, to measure systems humans built. Now they are, at last, tamper-resistant. The scores may finally reflect reality. What the reality reflects is, naturally, the next question.