OpenSSF Finding and Fixing Vulnerabilities Using AI

5.5 Triage

Most projects must perform triage; that is, they must prioritize reports.

Ideally, all vulnerabilities would be fixed immediately, but when there’s a lot to do, that’s often impractical. When handling everything immediately is impractical, findings must be prioritized so the most important vulnerabilities are addressed first. This is best done after validation, though if the validation process is overwhelmed, you may need to do some triage before validation.

Triage is fundamentally a risk decision. Risks, by definition, are important based on their:

  1. Likelihood. In particular, consider the preconditions necessary for the vulnerability to be exploitable and what is required for an attacker to exploit it [Yan2026]. A vulnerability that an unauthenticated attacker can exploit remotely is more likely to be targeted than one that requires a local authenticated user.
  2. Impact. A remote code execution (RCE) vulnerability is typically considered an extremely high impact because often an attacker can make the system do almost anything. In contrast, a vulnerability that can only reveal public or unimportant information would have low impact.

Most systems have some sort of “severity rating” classification system to help you identify the vulnerabilities most needing to be addressed [Rogers2025]. Unfortunately, many report that AI systems aren’t very good at estimating the severity of a vulnerability (they often report both too high and too low), so if you use an AI to estimate in a way that matters, have a human review the estimates.

You can make AI severity estimates more reliable by giving the AI your threat model and having it answer specific questions building on that before it assigns a severity.

Modern AI systems have become increasingly good at chaining many defects together into vulnerabilities. Something that appears unexploitable may, when combined with other defects, constitute a serious vulnerability. Thus, it’s reasonable to triage the “most dangerous vulnerabilities” first, but do not ignore other defects that appear unexploitable. Fix the other defects, as resources permit, so they don’t become part of a chain leading to an exploit.

Don’t assume an AI-found vulnerability will stay unknown for long. When many people use similar AI tools on the same code, they often find the same bugs around the same time. The Linux kernel security team reports that bugs found with AI assistance “systematically surface simultaneously across multiple researchers, often on the same day” [Linux-SecurityBugs].

A robot and woman categorize bugs into red (dangerous) and green (not dangerous); the woman overrides the robot’s categorization

Quiz

Q1. We warned that AI systems tend to be especially unreliable when estimating one particular aspect of a vulnerability report. What is it, and what do we recommend as a result?

  1. The exact file and line number of the defect; this material recommends always re-running static analysis tools to confirm locations.
  2. Severity ratings; since AI often rates severity too high or too low, a human should review any severity estimate that matters.
  3. Whether a finding is a duplicate; this material recommends discarding any finding that AI can’t confidently de-duplicate.
  4. The programming language used in the vulnerable component; this material recommends manually re-identifying the language before triage.
Show answer Answer: B