AI Struggles to Patch Vulns Without Adult Supervision
The Register, Thursday, August 6th, 2026
AI-generated security patches fully succeed only 26% of the time without human review.
Researchers at 1Password's Off-by-1 Labs tested ChatGPT 5.5 and Claude Opus 4.8 on six CVEs, generating 6,080 patches to evaluate effectiveness.
Only 26% fully resolved vulnerabilities without altering application behavior. Another 20.1% fixed issues but changed functionality, while 49.3% failed to remediate the exploit and 2.3% introduced new security problems.
The study concludes that autonomous LLM-driven patching poses significant long-term risks, since the cognitive burden of reviewing mostly incorrect patches can exceed the effort of manual remediation.