LIVE · cybersecurity feed
Live wire
Security Affairs newsletter Round 589 by Pierluigi Paganini – INTERNATIONAL EDITIONWebmail CSS Attacks Expose a New Risk for AI-Powered Email ToolsMetabase Zero-Day Exploited in the Wild, Exposing Admin Access and Sensitive DataCritical One-Click Vulnerability in Atlassian’s Rovo AI Exposed Enterprise DataCVE-2026-8037 · CISA Adds Progress LoadMaster Command Injection Flaw to KEV CatalogSensitive Info Goes Into ‘No Reply’ Emails Constantly. This Guy Sees It AllAtlassian Rovo Can Be Tricked Into Sending Jira and Confluence Data to AttackersNew CSS Attacks Can Break Webmail Defenses to Steal Passwords and TokensCVE-2023-38646 · Metabase Zero-Day Exploited in Wild Allows Admin Access Without AuthenticationCVE-2026-18577 · N-able Issues N-central Hotfix 2 as Attackers Reach Managed Systems and Persist
vulnerability

Three in four AI-generated vulnerability patches leave something broken

Ask a frontier model to patch a real vulnerability and it will hand you something that looks like a fix. It reads like the patch a maintainer would write. When there is a test, it often passes. Roughly one time in four, it is a fix. Researchers at 1Password graded 6,080 patches for six freshly disclosed CVEs, and the failures are rarely the obvious kind: an exploit path gated behind a check with t

zeroday.news ·

New research indicates that large language models (LLMs) tasked with patching software vulnerabilities frequently produce fixes that are incomplete, introduce new flaws, or alter expected program behavior. A study by Off-by-1 Labs, a security research group within 1Password, found that roughly three out of four AI-generated patches for real-world vulnerabilities left something broken.

The researchers evaluated 6,080 patches for six recently disclosed Common Vulnerabilities and Exposures (CVEs), using both ChatGPT 5.5 and Claude Opus 4.8. While many of the AI-generated patches appeared correct and often passed initial tests, closer inspection revealed significant issues.

A common failure mode was the "fragile fix," where models addressed a specific exploit path demonstrated in a reproducer but failed to resolve the underlying general problem. This left the vulnerable code intact and exploitable through alternative means. For instance, in CVE-2026-8512, a Chromium bug related to macOS folder change monitoring, both models routinely implemented only one of two necessary claims on a go-between object, leaving the flaw present but moved.

Approximately half of the evaluated patches failed to fully close the original vulnerability, leaving at least one exploitable path open. About one in twenty introduced entirely new vulnerabilities, sometimes in addition to not fixing the original issue. Other patches closed the original bug but inadvertently changed the software's behavior, such as rejecting previously accepted inputs.

The study highlighted a case with Freenginx, a web server, involving a use-after-free memory bug. An AI-generated fix for this flaw, submitted through the Patch the Planet initiative by Trail of Bits and OpenAI, was rejected by maintainers because it only repaired two of three vulnerable locations. Both the rejected AI patch and the maintainer-written fix also introduced a new way to crash the server. Off-by-1 Labs subsequently reported this new crash, which was fixed on July 2. In a separate case study, 270 attempts by ChatGPT 5.5 to patch the original Freenginx flaw, while 114 were judged to close the original hole, every one of those 114 introduced a new problem.

The Linux kernel privilege escalation vulnerability, "Copy Fail," disclosed in April, also presented challenges. The upstream fix for this bug involved reverting a memory optimization, which initially introduced an off-by-one heap write that required a subsequent correction. Roughly a third of the AI-generated patches for "Copy Fail" regenerated this flawed revert, including the off-by-one error. Additionally, a second, unrelated flaw in the same code section was overlooked by every AI patch, even those that edited the exact file.

The researchers noted that the models tend to fix only the specific bug described in the ticket, even when other vulnerabilities are present in the code they are modifying. This "patch the example, not the bug" approach was a recurring theme.

The quality of guidance provided to the LLMs significantly impacted their success. Prompts containing correct fix directions resulted in bug closure approximately two-thirds of the time. However, prompts offering plausible but incorrect directions drastically reduced success rates to about one in six. The study concluded that providing confidently wrong advice is more detrimental than offering no advice at all.

While an iterative patching process and richer correct context improved results, neither was as impactful as the accuracy of the initial guidance. The study also found that model performance varied wildly across different codebases; for example, Claude cleanly fixed an Exim remote code execution bug in about three-quarters of attempts but managed under one percent for a Gemini CLI trust-bypass advisory.

The cost of generating and validating each patch attempt ranged from two to three dollars, which is inexpensive compared to an engineer's time. However, the researchers emphasized that LLM-produced patches still necessitate review by a skilled engineer with domain expertise. This manual review is costly because understanding a patch well enough to certify its security implications requires at least as much effort as writing a known-good patch from scratch.

Due to the high volume of patches, automated validators were used, cross-checked against each other and spot-checked by humans. The automated reviewers did not catch all new vulnerabilities, such as the "Copy Fail" off-by-one error in many instances, suggesting that the study's reported new-vulnerability rate should be considered a floor. The tooling used in the study has been made public to allow organizations to run it against their own fixed bugs and derive repository-specific metrics.

vulnerabilitypatchai
ShareXLinkedInWhatsAppFacebook

More News

view all →
vulnerability

Week in review: Cisco fixes IMC bug, Patch Tuesday forecast, Black Hat USA 2026

Here’s an overview of some of last week’s most interesting news, articles, interviews and videos: Mapping the malware blast radius a single alert won’t show you In this interview with Help Net Security, Mike Wiacek, founder and CTO of Stairwell, explains Backstory, an AI agent that takes a single alert and works outward to map how far a malware campaign spread. He walks through the research behind

breach

Hackers breach TrueConf to trojanize client installers with backdoors

The Head Mare hacktivist group has been exploiting vulnerabilities in unpatched TrueConf video conferencing servers to replace client installers with malicious versions that deliver backdoors. [...]

phishing

U.S. Defense Manufacturer IEH Hit by Phishing Attack, Exposing Potentially Export-Controlled Data

Defense manufacturer IEH Corporation has disclosed a phishing attack that compromised an employee's Microsoft 365 inbox. The breach, discovered on August 4, potentially exposed sensitive export-controlled military data, customer information, and engineering documents. While no data exfiltration has been confirmed, the incident highlights the risks associated with sophisticated social engineering tactics targeting critical infrastructure suppliers.

malware

SECURITY AFFAIRS MALWARE NEWSLETTER ROUND 109

Security Affairs Malware newsletter includes a collection of the best articles and research on malware in the international landscape Malware Newsletter Fake Xeno Roblox Cheats Deliver Powerful Java Stealer Through Discord and Forums DarkSword’s Panel Sprawl: How One Body Hash Unravels a Six-Panel, Two-Codebase Operator Cluster Distributed npm Package Cluster Delivers Cross-Platform RAT Targeting

zero-dayhigh

Security Affairs newsletter Round 589 by Pierluigi Paganini – INTERNATIONAL EDITION

A new round of the weekly Security Affairs newsletter has arrived! Every week, the best security articles from Security Affairs are free in your email box. Enjoy a new round of the weekly SecurityAffairs newsletter, including international press. Palo Alto Networks Faces China Cybersecurity Review Amid Rising Tech Tensions Metabase Zero-Day Exploited in the Wild, […]

ransomware

Ransomware gangs skip the CEO, head straight for the 40-something IT manager

Gen Xers who feel triggered by this should remember to unplug the network cable and call the cops