Google DeepMind introduces double-blind AI evaluations with encryption

Google DeepMind has implemented a double-blind evaluation system for proprietary AI models, using a cryptographically protected environment to keep both the model and the evaluation questions mutually hidden. This pilot method, carried out in collaboration with the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons, represents a significant step towards more reliable and transparent evaluations of AI model capabilities.

Quick Answer

The new double-blind evaluation system by Google DeepMind uses encryption to protect both AI models and evaluation questions. This approach aims to prevent benchmark contamination, ensuring that results reflect the model's true capabilities. The technology used includes Google Cloud’s Confidential Computing, Intel TDX, and encrypted NVIDIA H100 GPUs.

The problem of benchmark contamination

Benchmark contamination is a growing issue in AI model testing. When a model or its developer has access to test questions before evaluation, a high score might simply reflect familiarity with the benchmark rather than the model's actual capabilities. Traditionally, safeguards like zero-logging policies and contractual restrictions have helped keep evaluation questions confidential, but cryptographic safeguards can add an extra layer of protection.

Technological architecture of double-blind evaluation

The system uses Google Cloud’s Confidential Computing technology to place the model and evaluation data within a protected environment. The evaluator cannot access Google’s model parameters, while Google cannot access the evaluator’s test questions. The pilot was run on a Google Cloud A3 Confidential VM, using Intel TDX host memory encryption and an NVIDIA H100 Confidential GPU. Hardware encryption and remote attestation were used to keep benchmark prompts and model parameters isolated while simultaneously verifying the software environment.

This approach is designed to reduce a long-standing compromise in external evaluations of AI models: in the past, evaluators had to choose between providing sensitive test material to model developers or asking companies to expose proprietary model parameters. Google DeepMind states that this method could be particularly useful for sensitive evaluations involving cybersecurity and government bodies.

Evaluation of Gemini 2.5 Flash Lite

The project involved AVERI, which evaluated Gemini 2.5 Flash Lite using confidential prompts from MLCommons’ AILuminate security benchmark. This benchmark covers critical areas such as cyberattacks, chemical and biological hazards, hate speech, self-harm, and the elicitation of violent crimes. Separately, the Singapore AI Safety Institute tested the model using confidential prompts focused on harmful content in the context of Singapore.

Limitations and future considerations

Despite the progress, the pilot leaves some questions unanswered. Notably, Gemini 2.5 Flash Lite scores or a detailed task-by-task analysis of the results were not published. The technical report acknowledges several limitations, including the inability to fully inspect some proprietary inference codes and the need to rely on Google for attestation verification. MLCommons also warned that technical secrecy alone is not sufficient; legal protections and careful benchmark management are necessary.

Implications for the AI industry

The greater significance of this experiment does not lie in the scores Gemini achieved on a security benchmark, but in the ability for AI companies to demonstrate that their benchmark results were obtained fairly, without influence from evaluators or developers. This distinction could become increasingly important as benchmark scores influence decisions by regulators, researchers, and companies.

For IT leaders evaluating vendor claims, this method could eventually provide stronger evidence that AI models were tested against independent, previously unseen benchmarks. However, until the process becomes independently reproducible and detailed results are released, buyers should still question who provided the benchmark, who evaluated the results, what conclusions were disclosed, and which parts of the system require trust in the model provider.

Towards an industry standard

For double-blind testing to become a significant industry standard, the process will need to be independently reproducible, transparent in methodology, and scalable across models and benchmarks. Without these elements, the industry could end up with safer tests without necessarily achieving more reliable results.

This development comes as Google is restricting the release of specialized models like Gemini 3.5 Flash Cyber, underscoring the importance for companies to examine independent performance proofs and test checks before adopting specialized AI models.

Implications for Data Security and Governance

The double-blind approach adopted by Google DeepMind raises important questions about the security and governance of sensitive data. The use of technologies like Intel TDX and NVIDIA H100 Confidential GPUs represents a significant step forward in protecting data during AI model evaluations. However, relying on Google for attestation verification introduces a level of trust that could be problematic for some highly sensitive applications.

The Role of Cloud Infrastructures

Google Cloud played a crucial role in this project, offering secure hardware and software infrastructures for conducting evaluations. The A3 Confidential VMs and remote attestation technologies allowed both model parameters and evaluation prompts to be kept isolated. This demonstrates how cloud infrastructures can be adapted to support secure evaluations of proprietary AI models.

Challenges for Reproducibility and Transparency

One of the main challenges highlighted by the technical report concerns the independent reproducibility of results. The lack of specific details on scores and reliance on Google for attestation verification limit the possibility of independent evaluation. To become an industry standard, the process will need to be open to third-party verification and transparent in its methodology.

Impact on Regulators and Regulations

The adoption of these double-blind evaluation techniques could significantly influence AI sector regulations. Regulators may require AI models to be evaluated according to independent and reproducible standards before being approved for public use or in critical sectors such as healthcare or defense. This approach could also influence companies' purchasing decisions, which may prefer models with transparent and independent evaluations.

Future Perspectives for AI Evaluations

The future of AI evaluations may involve more cryptographic techniques and hardware isolation to ensure the security and independence of tests. However, a balance between security and transparency will be necessary. Companies will need to develop protocols that allow rigorous evaluations without compromising the intellectual property or competitiveness of their models.

Considerations on Gemini 3.5 Flash Cyber

Google's announcement regarding the limited availability of Gemini 3.5 Flash Cyber underscores the importance of independent evaluations for specialized models. Companies will need to carefully examine performance proofs and test checks before adopting these models. The ability to demonstrate that benchmark results were obtained fairly could become a key factor in adoption decisions.

While Google DeepMind's pilot represents a significant step towards safer and more independent evaluations of AI models, many challenges remain. The industry will need to work to develop reproducible and transparent standards while ensuring the protection of intellectual property. Only with these advancements can double-blind evaluations become a reliable tool for the responsible adoption of AI in critical sectors.

Additional Resources

Editorial Note and Disclaimer

The guides and content published on GoYou are the result of independent research and analysis activities, for informational, educational, and in-depth purposes.

GoYou does not constitute a journalistic publication or an editorial product under Law No. 62/2001 and does not provide real-time information.

The GoYou project does not provide professional, technical, legal, or financial advice and disclaims all liability for the improper use of the information published.

In the Crypto sector, every investment involves risks: readers are invited to always inform themselves independently before making any decision.