Definition
The first agreed methodology, developed jointly across multiple governments, for how outside experts should test and measure AI system safety and capability — creating a common yardstick instead of every country or lab grading AI on its own terms. It's meant to make evaluations from different testers actually comparable to one another.
Why it matters
Boards relying on 'independent AI safety testing' claims should know whether that testing follows a recognized cross-government methodology like this one, or an ad hoc, unverifiable process.