A vulnerability is a defect that an attacker can exploit to violate some security requirement. By contrast, a finding is only a potential vulnerability. Not all findings from AIs (or humans) are real vulnerabilities [Ottenheimer2026-05-26]. Thus, once you have a finding, you need to validate it to independently determine whether it is a vulnerability, a defect but not a vulnerability, or neither.
Even when a finding isn’t a defect, it might suggest improvements for the future. However, here we’ll focus on validating if a finding is a vulnerability.
5.4.1 Defining what a vulnerability is
A key to validating findings is having a threat model and/or security requirements. Otherwise, there’s no way to tell whether something is security-related, making it difficult to determine whether it’s a vulnerability. It’s also vital to have a validation infrastructure, that is, a way to run tests to confirm or refute claims [Ziegler2026].
An important decision is whether defense-in-depth countermeasures should be disabled or ignored during validation. An example of a countermeasure is its own process sandboxing, which is different from the sandbox around the AI used during development. A mechanism is only a defense-in-depth measure if it’s an additional security measure, not a required one. Thus, from a security perspective, it’s better to consider something a vulnerability even if another countermeasure would have prevented it. For example, the Mozilla Firefox project considers a finding as a vulnerability even if a sandbox would have prevented its exploitation. Their rationale is that “Real-world attackers generally need to chain multiple exploits together to escalate privileges through one or more layers [of mitigations like sandboxing] [Grinstead2026-05]. Of course, vulnerabilities that can penetrate existing countermeasures should be prioritized.
5.4.2 Defining the proof required
Projects must define the level of proof required for a finding to be considered a vulnerability within their project. Here are some key considerations:
The gold standard for validating a finding is to generate a PoC, that is, a demonstration using a specific input to the program that an adversary could supply that causes the system to fail its security requirements.
In practice, many projects consider being able to trigger abnormal execution (such as a crash, memory error, or sanitizer violation) sufficient to prove that a discovered defect is real and reachable. For example, Firefox classifies findings as high security vulnerabilities “based on predictable crash symptoms such as use-after-free or out-of-bounds memory issues being reported by AddressSanitizer, and our threat model assumes that any of them could be exploitable with sufficient effort. This reduces the risk of a false negative during exploitability analysis, and more importantly it allows us to focus our resources on finding and fixing more vulnerabilities.” [Grinstead2026-05]
Some AI systems incorporate guardrails that prevent the creation of a PoC (since it can also be used for an attack), making them useless for creating PoCs.
A project may decide that more general information describing how to exploit the vulnerability, without details, is adequate evidence.
In the long term, it’s often worthwhile to fix defects even if a PoC isn’t found, since, even if a PoC isn’t found, a different AI system might find a way to exploit it. Typically, later triage focuses on findings known to be exploitable first.
It helps to say explicitly what level of evidence you require. From weakest to strongest, a PoC may show that: an input triggers abnormal behavior, such as a crash or a sanitizer error; the input violates a specific security property, as shown by a test that checks for that violation; an attacker can gain a specific capability against a realistically deployed system (e.g., code execution, reading or writing files, or bypassing authentication), possibly by chaining several defects; or an attacker can do so reliably, that is, there’s a working exploit. Some researchers use separate terms for these levels. For example, Antiproof defines a “proof of vulnerability” as “an executable test that uses a bug oracle, such as a test harness with assertions, to demonstrate a violation of a security property in isolation”, and a “proof of exploitability” as “an end-to-end test that demonstrates an attacker capability against a deployed system” [Shakevsky2026]. As noted above, Firefox treats even the first level (a sanitizer-detected memory error) as enough to call a finding a high-severity vulnerability. Whatever level you choose, the inverse doesn’t hold: “failure to produce a working PoC is not proof of a false positive” [Yan2026].
Part of validating findings typically involves using AI to review the proposed findings. In principle, a stronger model will be better at this, but it’s no guarantee. [Ziegler2026] reported that when using Mythos Preview, its “judgment results were more mixed than its discovery results… It rejected false positives better than many predecessors, but sometimes lost true positives when evidence did not formally satisfy its criteria or when the intended rule was broader than the written one.”
5.4.3 Tips for validation
Here are some tips for validating that a finding is (or is not) a vulnerability.
There’s no magic to validation. Validating claims has long been a part of software development. In validating AI findings, [Kholoosi2025] reported that “Practitioners in our study consistently emphasized a layered validation approach involving manual inspection, sandbox testing, peer review, and cross-checking against established standards such as OWASP or NIST.” It’s all part of using multiple stages to validate (and deduplicate) a report [Rogers2025].
The verifier agent should be independent from the discovery agent. “A useful way to frame this is that validation involves independent reproduction rather than self-review. Asking the discovery agent to reconsider its own conclusion can preserve the same context, assumptions, and reasoning errors.”
The verifier should instead independently attempt to disprove the finding and produce reproducible evidence for its conclusion. [Yan2026] says, “Run the verifier in a fresh container without a shared filesystem or conversation history. If the verifier is exposed to the discovery agent’s reasoning, it may simply agree instead of testing the claim. Thus, give the verifier only (1) the proof of concept or written finding and (2) the codebase, so it can search for mitigations the finder missed (e.g., upstream validation, auth gates, type constraints, or unreachable code).”
Another approach is to “prompt the verification agent to disprove the discovery agent’s findings [while still providing the full set of information available]. Have the verifier assume each finding is a false positive and search for reasons the finding is wrong. Include clear criteria that the verifier agent can use to determine if the finding is a true positive. This [approach] matters most when the discovery agent’s output doesn’t include a PoC. Aim to exclude as many non-exploitable findings as possible to reduce effort on manual reviews.”
“Across the teams we’ve worked with, adding an adversarial verifier roughly halved the rate of non-exploitable findings from the discovery phase. Requiring that verifier to also build a proof of concept confirming the exploit brought the false positive rate to near zero. Together, these two steps helped to reduce the downstream triage and patching load significantly.”
“One team scanning open-source packages built a verification step that helped to close the loop: scan the package, generate a proof of concept, then deploy a mock application that uses the package and triggers the PoC. Their take was that: ‘Validation is the biggest holdup and the PoC is the validation.’”
Do not treat agreement among AI reviewers as proof. In a 2026 case study of an adversarial multi-agent review process, “ten dedicated [AI] reviewers unanimously endorsed a non-existent Bleichenbacher padding oracle in OpenSSL’s CMS module; it was killed only by a single empirical test” [Agarwal2026]. Here, empirical confirmation means that the target software was actually run with a test, a PoC, or something equivalent that demonstrated the claimed failure in the target itself. The author concluded that “empirical verification, not consensus count, is what changes our belief”, and added a mandatory gate: “no candidate reaches disclosure without empirical confirmation.” With that gate, the pipeline “killed roughly 79% of 171 candidates before advancing to disclosure” while still producing 4 CVEs and 8 merged security-related fixes.
PoCs must measure target behaviour, not their own artifacts. A PoC often includes a script or program that supplies the input and then checks the result, and that check can be wrong. In that same study, one proof-of-concept program “detected its own nonce computation rather than the library’s leak”, falsely confirming the vulnerability [Agarwal2026].
A PoC needs checking too, and the fix must cause the PoC to fail. A 2026 reproducibility study of published LLM/agent-driven vulnerability research reran the researchers’ own verification checks. In 20 of 30 cases the check still reported the vulnerability on the patched version, and 7 of 19 negative controls “still trigger on benign input”. As the study says, “A trigger on the vulnerable build is not evidence of CVE-specific reproduction without a clean patched counterfactual.” [Chen2026]
So when you accept a PoC as evidence, confirm that it:
triggers the specific claimed failure (not just any crash) on the vulnerable version,
does not trigger on benign inputs, and
does not trigger once the vulnerability is fixed (check this when verifying the fix).
Ask a separate AI instance to disprove the findings, and if it can’t, demand a PoC. Each step can dramatically reduce false positives, and doing them as separate steps means the AI system won’t abandon findings too quickly.
Quiz
Q1. Why does the Mozilla Firefox project consider a defect to be a vulnerability even when an existing countermeasure, such as a sandbox, would have prevented exploitation?
Because a defense-in-depth measure is by definition an additional protection, not a required one; relying on it reduces the layers an attacker must defeat.
Because Firefox’s threat model assumes that end users will eventually disable all sandboxes, so no countermeasure can be trusted to remain active.
Because the CRA legally requires vendors to report every defect regardless of whether it’s exploitable, so classification is a required compliance formality.
Because AI systems can’t reliably detect which countermeasures are active, Firefox treats every defect as maximally severe to be safe.