OpenSSF Finding and Fixing Vulnerabilities Using AI

5.7 Verify fixes

After generating candidate fixes, have AI and then humans review them to verify them.

This is useful for a fix by a human or by an AI. However, it’s especially important for AI-created code. Even very good AI writes vulnerable code. In particular, an AI can be good at finding vulnerabilities yet still write code with lots of vulnerabilities. This might be surprising, but remember, AI is trained on a large amount of insecure software. It’s more difficult to get an AI to do something contrary to its training dataset [CSA2026] [Zimmer2026]. Modern AI is fundamentally probabilistic, so it can be difficult to predict exactly what it will generate for a given request [Kholoosi2025].

5.7.1 Have AI review the fix first

The first step is to have an AI review the proposed fix for the vulnerability. “Have a new discovery agent probe the patch as an attacker to confirm the patch is comprehensive” [Yan2026]. What can be especially helpful is to have AI write tests for the fixes. For example, [Chrome2026] reports that they use test-writing agents to “help write tests for fixes. These agents can ensure that tests work across supported platforms and configurations before a developer reviews the fix, saving up to weeks of developer time.”

Run the validated PoC against the fixed version and confirm that it no longer succeeds. Then have a separate agent try variations of the PoC, since a fix that blocks only the exact PoC input isn’t a full fix.

5.7.2 Have humans review the fix

While AI review is a good first pass for proposed vulnerability fixes, they still require expert human review. As we noted earlier, AI-generated fixes often are incomplete, add vulnerabilities, or break functionality.

Some problems with initial AI-proposed fixes are especially common, and humans should especially look out for them:

Even if an AI’s initial fix is wrong, that doesn’t make its proposed fix useless. One reporter found AI fixes “most useful for simply understanding what the problem actually was in the code – sometimes I didn’t understand the issue from the description, but the suggested fix revealed to me what was wrong, and what would (could?) fix it.” [Rogers2025] Sometimes the proposed fixes shouldn’t be the final fix, but they can still provide guidance to help developers find a correct solution.

5.7.3 Apply your usual verification processes

Be sure to apply all of your usual verification processes. That includes your peer review gates, CI/CD pipelines, and branch protection rules. “Leveraging organizational safeguards such as peer-review gates and branch protection rules can help mitigate potential individual complacency regarding AI-generated security suggestions” [Kholoosi2025].

Technology can help. However, as noted in [Anthropic2026-05], now “the bottleneck in fixing bugs like these is the human capacity to triage, report, and design and deploy patches for them.”

5.7.4 Turn findings into durable checks

AI analysis is non-deterministic; don’t depend on rerunning it to catch the same problem again. Instead, turn what you learned into deterministic checks that run on every change, as part of the change.

At the least, modify the test suite to add the regression test you created earlier to verify the problem. You may also want to add:

  1. A static analysis rule that detects the vulnerable pattern (e.g., a Semgrep rule or CodeQL query).
  2. A fuzz harness for the code involved, if it’s the kind of code fuzzing handles well (e.g., parsers).
  3. A hardening or API change that makes the whole class of vulnerability impossible or unlikely (see the upcoming section on hardening).

AI can often help implement all of these. The general principle is that if something is important, enforce it with deterministic mechanisms instead of relying on an AI to notice it every time.

Q1. You’ve applied a proposed fix, and the validated proof of concept (PoC) no longer succeeds against the fixed code. What should you do next?

  1. Close the finding, since the PoC no longer succeeding shows the vulnerability is fixed.
  2. Rerun the AI that found the vulnerability, and treat a clean result as proof the fix is complete.
  3. Have a separate agent try variations of the PoC, since a fix might defeat only the exact PoC input rather than the underlying vulnerability.
  4. Remove the PoC from the test suite so it doesn’t reveal exploit details.
Show answer Answer: C