What happened
Anthropic's periodic Risk Report (published every 3-6 months, this edition running 186 pages) assesses catastrophic risks from its frontier models across four threat categories: misalignment in high-stakes settings, automated AI R&D, chemical/biological weapons uplift, and cross-cutting acceleration dynamics. The company rates overall risk from automated R&D as 'low' but notes 'meaningful acceleration starting in early-to-mid 2025, though by less than a factor of 2' and says it is less confident than in prior reports because some task-based evaluations have saturated. On its internal CoBench evaluation, Claude Mythos 5 scored 50.3% and Model 2 scored 62.8%, versus an estimated 85% threshold for full substitution of Anthropic's own research staff. The report also raised its misalignment risk rating from 'very low' to 'low' due to greater uncertainty following recent model-behavior incidents, and documents redaction of sections covering autonomous replication/adaptation capabilities.
Why it matters
This is a frontier lab's own quantitative disclosure of how close its models are to crossing capability thresholds that trigger stricter safety protocols (ASL-3) — the CoBench gap (50-63% vs. 85% threshold) and the acknowledged 'early signs of acceleration' give boards and CISOs a concrete, lab-published yardstick for how fast automated AI R&D risk is actually moving, distinct from vendor marketing claims.
Action needed
Brief the board and AI risk committee on the CoBench acceleration metrics and monitor Anthropic's threshold updates as an external benchmark for internal AI-agent autonomy governance policies.