After generating candidate fixes, have AI and then humans review them to verify them.
This is useful for a fix by a human or by an AI. However, it’s especially important for AI-created code. Even very good AI writes vulnerable code. In particular, an AI can be good at finding vulnerabilities yet still write code with lots of vulnerabilities. This might be surprising, but remember, AI is trained on a large amount of insecure software. It’s more difficult to get an AI to do something contrary to its training dataset [CSA2026] [Zimmer2026]. Modern AI is fundamentally probabilistic, so it can be difficult to predict exactly what it will generate for a given request [Kholoosi2025].
The first step is to have an AI review the proposed fix for the vulnerability. “Have a new discovery agent probe the patch as an attacker to confirm the patch is comprehensive” [Yan2026]. What can be especially helpful is to have AI write tests for the fixes. For example, [Chrome2026] reports that they use test-writing agents to “help write tests for fixes. These agents can ensure that tests work across supported platforms and configurations before a developer reviews the fix, saving up to weeks of developer time.”
Run the validated PoC against the fixed version and confirm that it no longer succeeds. Then have a separate agent try variations of the PoC, since a fix that blocks only the exact PoC input isn’t a full fix.
While AI review is a good first pass for proposed vulnerability fixes, they still require expert human review. As we noted earlier, AI-generated fixes often are incomplete, add vulnerabilities, or break functionality.
Some problems with initial AI-proposed fixes are especially common, and humans should especially look out for them:
Even if an AI’s initial fix is wrong, that doesn’t make its proposed fix useless. One reporter found AI fixes “most useful for simply understanding what the problem actually was in the code – sometimes I didn’t understand the issue from the description, but the suggested fix revealed to me what was wrong, and what would (could?) fix it.” [Rogers2025] Sometimes the proposed fixes shouldn’t be the final fix, but they can still provide guidance to help developers find a correct solution.
Be sure to apply all of your usual verification processes. That includes your peer review gates, CI/CD pipelines, and branch protection rules. “Leveraging organizational safeguards such as peer-review gates and branch protection rules can help mitigate potential individual complacency regarding AI-generated security suggestions” [Kholoosi2025].
Technology can help. However, as noted in [Anthropic2026-05], now “the bottleneck in fixing bugs like these is the human capacity to triage, report, and design and deploy patches for them.”
AI analysis is non-deterministic; don’t depend on rerunning it to catch the same problem again. Instead, turn what you learned into deterministic checks that run on every change, as part of the change.
At the least, modify the test suite to add the regression test you created earlier to verify the problem. You may also want to add:
AI can often help implement all of these. The general principle is that if something is important, enforce it with deterministic mechanisms instead of relying on an AI to notice it every time.
Q1. You’ve applied a proposed fix, and the validated proof of concept (PoC) no longer succeeds against the fixed code. What should you do next?