Every business that deploys AI will want to know: does this system perform reliably and safely for my use case and with my data? But it’s not as simple as hooking everything together and running a test.
Deployers, like banks, need to keep sensitive data safe. AI solution providers, like frontier labs, are very protective of model weights. And most critically from our perspective, benchmark providers must have both technical and legal mitigations in place to ensure the models being tested won’t compromise the integrity of their benchmarks by learning from them. Importantly, secrecy of the evaluation itself is not sufficient. A robust benchmark stewardship program is also a necessary component of operating a high-integrity benchmark over time.
Demonstrating best-in-class benchmark integrity is why we participated in this first-ever, double-blind evaluation proof of concept with Google DeepMind, OpenMined, and AVERI. This approach keeps every party protected by design, rather than relying solely on legal contracts, as part of the first-ever double-blind evaluation of a closed-weight model. Through this method, we have cryptographic guarantees that our evaluation components are able to be used without contaminating the benchmark or compromising its integrity.
Our approach to benchmark integrity is what made AILuminate™ the obvious choice for this proof of concept. We provided a reserved subset of the AILuminate™ safety benchmark prompts to be used in running this evaluation, ensuring that no GDM model had ever been exposed to this specific evaluation set before. AVERI used OpenMined’s secure computation to run these prompts on a containerized instance of a Google DeepMind model, with cryptographically provable protections for both the model’s weights and our evaluation.
In this proof of concept, AILuminate characterized an AI system’s reliability without exposing our test data to the developer, and without exposing the developer’s proprietary model weights to the instrument maker (MLCommons) or the auditor (AVERI).
We are eager to see this type of IP-safe, saturation-resistant testing become the standard for applied tests of AI models. In support of it, we’ve already done this sort of work on the most sensitive of data: health informatics. Through our MedPerf program, we’ve been building Trusted Execution Environments (TEEs) just like these, which enable multiple parties to bring encrypted data together in one place, run high-integrity evaluations, and get results, all without seeing any other party’s sensitive information.
We’re working toward a near future where AI adopters can quickly and easily run trustworthy evaluations of AI systems for their purpose-specific deployments. An environment that structurally guarantees the confidentiality of model, data, and test paired with the assurance of never-before-seen prompts from an independent organization, provides the clearest path toward informed decision making about AI adoption, deployment, and reliable operation.
MLCommons is working toward a near future where AI adopters can quickly and easily run trustworthy evaluations for their purpose-specific deployments. We invite practitioners, researchers, and deployers to review the technical details and explore opportunities to engage with the AILuminate framework and the Benchmark Stewardship Program at https://mlcommons.org/get-involved/.