toolcall.
ResearchAug 30, 2026, 20:26 UTC

Google DeepMind tests double-blind AI evaluations for Gemini

The pilot keeps benchmark prompts and model weights hidden from each other, aiming to make frontier model safety tests harder to game.

Google DeepMind illustration for double-blind AI evaluations

Google DeepMind says it has run a double-blind evaluation pilot for a proprietary Gemini Flash Lite model, using a secure environment designed to keep both benchmark prompts and model weights protected.

The test was built with the Singapore AI Safety Institute, OpenMined, AVERI and MLCommons. In the setup, external evaluation prompts are confined inside a cryptographic environment, while the model is also protected from the benchmark provider and auditor. The goal is to reduce benchmark contamination: the risk that a model has already seen test questions during training or evaluation preparation.

MLCommons said the pilot used a reserved subset of its AILuminate safety benchmark prompts that Google DeepMind models had not previously seen. AVERI ran the prompts with OpenMined secure computation on a containerized Google DeepMind model, producing reliability signals without exposing the test set to the developer or the model weights to the evaluators.

This is still a proof of concept, not a new public Gemini score. But it matters because AI labs, regulators and enterprises increasingly depend on external evaluations to judge model safety and capability. If double-blind testing becomes easier to run, benchmark results should become harder to inflate and more useful for real deployment decisions.

Sources

Mentioned

ai-safetyai-transparencybenchmarksgeminigoogle-deepmind