Google Deepmind wants to use a cryptographic method to stop ai models from seeing test questions in advance. A pilot project with the Singapore AI Safety Institute and other partners runs a double-blind test on a Gemini model.
If a test-taker knows the questions ahead of time, even a perfect score is worthless. Google Deepmind uses this image to describe a core problem in evaluating ai models: benchmark contamination. If a model has already seen the test questions during training, you can only trust the results so far.
To prevent that, Google Deepmind says it's launching the first double-blind evaluation of a proprietary frontier ai model. External tests stay locked in a cryptographic "box," so a model can't later use those questions to optimize itself specifically for the test.
For the pilot, Google is testing a model from the Gemini Flash Lite line against confidential benchmarks.
Highly sensitive external evaluations used to require a compromise, Deepmind says. Either the evaluators handed over their test prompts, which let the model provider see the questions in advance. Or the provider handed over its model weights and risked its intellectual property. A recent example of this dilemma was the delayed evaluation for theARC-AGI benchmark of Anthropic's fable 5, since the AI company enforces a 30-day data retention policy for its strongest models.
The double-blind evaluation is meant to eliminate that tradeoff. Google uses Confidential Space from Google Cloud'sconfidential computing portfolioto do it. The setup cryptographically verifies that both the external test data and the model stay private to their respective owners. The evaluator never sees the Gemini weights, and Google never sees the test prompts.
Until now, zero-logging protocols and contractual safeguards kept external prompts confidential. Adding technical and cryptographic protection is a big step forward for secure model evaluation, the company says.
The cryptographic proof aims to prevent contamination and protect sensitive data. Deepmind says this matters most for highly sensitive evaluations, such as cybersecurity or tests run by government agencies. Independent organizations could rigorously test advanced models without giving up data sovereignty or security.
Google hopes the effort sets a new standard for model oversight and helps the industry build more reliable and widely trusted AI systems. Google lays out the details on methodology and results in atechnical report.
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.