Skip to content
← Back to Insights

DeepMind Locks Model and Test Together: AI Evaluation Gets Its First Double-Blind Pilot

GenAI News AI Safety Gemini

TL;DR

Google DeepMind and independent evaluators used confidential computing to isolate Gemini weights from private test prompts, addressing a persistent conflict in AI benchmarking.

DeepMind Locks Model and Test Together: AI Evaluation Gets Its First Double-Blind Pilot

This pilot will change procurement decisions only under a testable condition: within six months, at least one other model developer must accept a comparable double-blind evaluation and publish a reproducible method. If Google DeepMind remains the only company to demonstrate the setup, buyers still cannot use it to compare vendors. On 2026-08-27, Google DeepMind announced what it describes as the first double-blind evaluation of a proprietary frontier model. The model tested was Gemini 2.5 Flash Lite.

Public leaderboards have a structural information problem. When an evaluator sends private questions through a model provider’s API, the provider can technically gain access to those questions. If the provider instead hands model weights to an external evaluator, it exposes intellectual property produced with substantial training investment. Contracts, zero-logging promises, and access controls constrain organizational behavior, but they do not prevent either side from reading the other side’s assets at the hardware layer. A separate contamination problem arises when public benchmark questions have already entered training data: a score can then reflect recall instead of performance on an unfamiliar task.

Google DeepMind ran the pilot with Singapore’s AI Safety Institute, OpenMined, AVERI, and MLCommons. The team loaded model weights and inference code into a Google Cloud A3 Confidential VM. It then brought reserved prompts from the AILuminate safety benchmark and a Singapore-focused harmful-content set into the same isolated environment. Host memory was protected by Intel TDX, while acceleration came from one 80 GB NVIDIA H100 Confidential GPU. Remote attestation checked the software environment before data arrived over encrypted connections. The model provider could not inspect the prompts, and the evaluators could not retrieve the weights.

The architecture protects information during execution; it does not certify that the benchmark itself is good. MLCommons still has to manage the representativeness, quality, and lifecycle of its prompts. The hardware, cloud attestation service, and deployment code also become part of a new trust boundary. Most importantly, neither Google DeepMind nor the independent report published Gemini 2.5 Flash Lite’s aggregate score or task-level results. The pilot therefore provides no basis for concluding that the model is safer or for ranking it against competitors.

Deployment constraints narrow the immediate use case. The demonstrated system fits the model onto a single NVIDIA H100, whereas larger frontier systems generally span multiple accelerators. This material does not quantify the performance overhead, operating cost, or security properties of a multi-GPU confidential-computing setup. An enterprise buyer can still use the method as a better due-diligence checklist: who controls the prompts, who scores the outputs, what code remote attestation verifies, and which findings are disclosed. A leaderboard number alone answers none of those questions.

Two observable developments matter over the next three to six months. Other model companies would need to adopt the protocol, and an independently validated multi-GPU version would need to appear. Without either development, the work remains a single-GPU engineering proof. If an outside laboratory can reproduce the process and release results, AI evaluation gains an auditable chain of evidence that contractual promises alone cannot provide.

Sources:

Get the latest insights

Join the newsletter to receive my latest articles on GenAI, AI Agents, and architecture.

No spam. Unsubscribe anytime.