What happened
1Password's new Off-by-1 Labs research team published (Aug 6, 2026) findings from 6,080 AI-generated patches across six recent CVEs: 53.9% failed to fully remediate the vulnerability or introduced new ones, while only 26% cleanly resolved the flaw without behavioral side effects. The team open-sourced its FLAWED tooling and datasets alongside the paper.
Why it matters
Provides hard evidence that autonomous AI patch-generation (e.g., OpenAI's Patch the Planet, Anthropic's Project Glasswing) requires expert human review before production deployment, directly informing how enterprises should gate AI-driven vulnerability remediation pipelines.
Applicability
AppSec teams adopting AI-assisted patching workflows should incorporate human review gates and consult the FLAWED benchmark before trusting autonomous patch generation at scale.