Definition
A way of testing AI models for safety or capability where neither the testers nor the AI company can see each other's confidential details, using a special locked-down computing environment (a 'secure enclave') to run the test blind on both sides. This prevents companies from tailoring their model to a known test, and protects testers' proprietary evaluation questions.
Why it matters
As governments move toward mandatory third-party audits of powerful AI models, a trustworthy blind-testing method is what stops those audits from being gamed or leaked.