Solutions  ·  2026-08-29

UK AI Security Institute open-sources optstop for compute-efficient frontier model evaluations

SolutionsMedium impactGlobal
AISI released optstop (Aug 27, 2026), an open-source Python package integrated with its Inspect evaluation framework that adaptively stops sampling once model-performance estimates meet precision/stabilization thresholds, reporting 57-97% trial savings across binary/ordinal/continuous benchmarks (MATH, GPQA Diamond, WritingBench) without shifting score estimates.
Frontier safety/capability evaluations increasingly require hundreds of millions of tokens; a government-backed, auditable early-stopping method materially lowers the cost of running rigorous dangerous-capability and safety evals, potentially increasing how often and how thoroughly labs and third-party evaluators test frontier models.
AI labs, third-party evaluators, and model-risk teams running Inspect-based evaluations should pilot optstop in shadow mode before enabling live stopping in safety-critical assessments.
UK AI Security InstituteGitHub
See this in the live feed Explore related AI security and governance findings — updated every morning.
Open the feed →