Do not presume that AI systems can always correctly fix all vulnerabilities.
It’s challenging to determine exactly how good AI systems are at creating fixes that fully fix the vulnerability, don’t introduce other vulnerabilities, and don’t interfere with system functionality. As of 2026, there’s evidence that even the best AI systems can struggle, especially when trying to correctly fix complex vulnerabilities.
![]()
A study by 1Password found that when the fix “had to touch multiple files, functions, or code paths, and introduce non-trivial changes” an AI would succeed only 26.0% of the time at generating a fix that fully resolved the vulnerability without materially changing application behavior. AI systems did not resolve the vulnerability, added a new vulnerability, or both, on average, 53.9% of the time [Hoodlet2026] [Mierczuk2026].
Indeed, they argue that, for complex vulnerabilities, using AI to fix them may be a poor use of resources as of 2026. Their report found that so few fixes worked correctly that “human auditors of LLM-generated patches are likely to spend the majority of their time reviewing and ultimately rejecting an avalanche of unnecessary code… the level of understanding one must build to confidently evaluate the full correctness of a vulnerability patch is often at least what would have been sufficient for a human programmer to produce a single, known-good patch in the first place” [Mierczuk2026].
Others have raised issues about that study or the way it was reported. That study only examined complex vulnerabilities, a key constraint the study clearly stated but that news summaries often omitted. In addition, a subsequent review found that some prompts told agents to apply the wrong fix or prohibited testing, and that only default reasoning settings were used (not the highest available setting). This review found that 86% of patches (2,634 of 3,067) blocked the supplied exploit, excluding cases where the upstream fix appears to have been used [Naik2026]. However, as that review clearly states, that upper figure is far too generous; simply blocking an exploit is not a full repair.
An in-depth 2026 analysis via “PatchBench” provides additional figures and examines a variety of vulnerabilities in C/C++ code. Once they countered memorization of previous fixes, and took steps to validate that the fix actually fixed the vulnerability (not merely an example input) and did not interfere with functionality, AI systems managed to correctly fix the vulnerability “roughly half” the time. The best agent could only solve 59% of the tasks in PatchBench, and 31% of the tasks (67 out of 213) couldn’t be solved by any of them [Shen2026-09].
It may seem surprising that AI systems can struggle to produce an accurate fix. AI systems can indeed generate a lot of code, especially simple code similar to what many others have done. This has misled some people to believe that AI can easily generate any kind of code. However, in many ways AI systems act like junior developers. AI can generate a lot of “straightforward” code, and that’s great because a lot of code is straightforward. However, AI is not good at applying specific local context or at complex situations. AI also has a propensity to generate insecure code, since it was trained on lots of insecure code. Completely fixing complex vulnerabilities without creating new ones or breaking functionality is where today’s AI systems are weakest.
Of course, AI can help generate many fixes. Many vulnerabilities are relatively easy to fix, where the fix is localized to a specific line or set of adjacent lines. AI can be especially good at fixing these once it’s told what needs fixing. AI can sometimes successfully fix more complex vulnerabilities, too. Perhaps most importantly, this information only applies to AI models as of 2026. We do not know how much better the AI models and underlying tools will become. That said, this limitation is unlikely to disappear instantly, and not everyone can use the best available systems.
As a result, it’s important to be aware that AI systems don’t always generate complete fixes. They may fix a special case but not the full vulnerability; they may introduce new vulnerabilities; and their fixes may interfere with correct operation. There’s some evidence that they especially struggle with more complex vulnerabilities. Plan accordingly.
Q1. What does the material conclude about AI-generated fixes for vulnerabilities, especially complex ones?