What happened
The International Network for Advanced AI Measurement, Evaluation and Science (formerly the International Network of AI Safety Institutes) published its first best-practice guidance document for third-party AI evaluators, completed at a Seoul meeting on the margins of ICML 2026. The document 'details: the importance and process of carefully defining evaluation objectives and selecting appropriate benchmarks; how to ensure comparability between evaluations...; how to iterate on capability elicitation...; and the processes and issues to work through in conducting evaluations and tracking results.' It builds on and complements NIST AI 800-2, targeting the growing ecosystem of independent AI evaluators across 10 member institutes. Companion pieces from Singapore's and Canada's AI Safety Institutes address system-level testing and evaluator information-sharing/selective disclosure norms respectively.
Why it matters
This is the first cross-government agreed methodology standard for third-party AI evaluation, directly shaping how regulators, auditors, and enterprises will benchmark and compare frontier model risk assessments going forward — a reference point for any board or compliance team building AI assurance programs.
Action needed
Map internal/vendor AI evaluation and assurance practices against the Network's best-practice recommendations, particularly on capability elicitation and comparability.