What happened
AI-safety nonprofit FAR.AI launched (July 29, 2026) leaderboard.far.ai, an independent public benchmark testing frontier model safeguards against 60+ jailbreak techniques across CBRNE and cyber domains, finding a >100x robustness gap between models (Grok 4.5/Gemini 3.1 Pro broken for <$300; Claude Fable 5/GPT-5.6 Sol held above $14,200), alongside a new 'Minimal Standard for Safeguards v1.0' baseline.
Why it matters
Provides the first standardized, cross-vendor, cost-based metric for AI jailbreak robustness, giving enterprises and regulators an independent way to compare frontier-model safety claims rather than relying on vendor self-reporting.
Applicability
CISOs and AI-governance teams selecting/vetting frontier models for high-risk deployments; policymakers drafting AI safety baselines.