What happened
Apollo Research publishes a framework for embedded evaluations of frontier AI developers, arguing that final-checkpoint testing cannot catch loss-of-control risks that arise during training and internal deployment (citing the OpenAI–Hugging Face incident as proof). The paper proposes a claim-based design in which evaluators with employee-equivalent access verify or falsify developers' safety claims, and defines four criteria any effective embedded evaluation must meet: 'Reduces risk, Informs the public, Good incentives, Fair to both sides.' It argues evaluators need ongoing access to all training data, rollouts, checkpoints, and internal deployment details, should publish by default with meta-transparency on redactions, and should disclose serious findings (e.g., a model attempting to copy its weights to external servers) within a short pre-agreed window. The framework builds on established practice for independent auditing in other high-stakes industries and is intended as a minimal starting point for developers and evaluators to adopt.
Why it matters
Embedded evaluation is the mechanism frontier labs just signed on to in the September 2026 pacing commitments, and this is one of the first concrete, implementable designs for what 'employee-like access' should actually deliver — directly relevant to any organization negotiating third-party AI assurance or preparing for the emerging embedded-evaluator regime.
Action needed
Map the four criteria and the claim-based access requirements against any third-party AI assessment clauses you are writing or signing, so access provisions match what independent evaluators argue is necessary.