HeadlinesBriefing favicon HeadlinesBriefing.com

Double-Blind AI Evaluations

Google DeepMind Blog •
×

AI benchmark results can be skewed if models see test questions in advance, a problem known as benchmark contamination. To address this, Google DeepMind is piloting the world's first double-blind evaluation of a proprietary frontier model. This approach keeps external evaluations in a cryptographic 'box' so they can't be used to optimize performance ahead of testing.

Google is partnering with the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons to test a Gemini Flash Lite model against confidential benchmarks in a privacy-preserving environment. This increases evaluation integrity by preventing the model from 'peeking' at the questions.

Previously, external partners used zero-logging protocols and contractual safeguards to keep test prompts confidential. Adding technical and cryptographic safeguards is a major step forward. Policymakers, researchers, and enterprises need trustworthy benchmarks to accurately reflect a model's true capabilities and safety.