The number of relevant tools and services has exploded. We can’t list them all. Here are pointers to specific ones you may find especially useful for this purpose, at least to give you an idea of what’s available.
First, we must acknowledge a problem: it can be challenging to find the tools and services to identify vulnerabilities with AI. There are so many blog posts, papers, and products related to AI and vulnerabilities that it can be difficult to find actual products, services, or systems for finding vulnerabilities in software [Bressers2025] [Rogers2025].
It’s also hard to evaluate them. Many published results come from the tool’s own developers, use different codebases and definitions of success, and report findings rather than validated vulnerabilities.
Be skeptical of benchmark numbers, especially for tools that classify isolated functions as vulnerable or not. Many such benchmarks have inaccurate labels: when researchers manually checked samples of functions labeled “vulnerable” in popular automatically labeled datasets, only 25% to 60% were labeled correctly [Ding2025]. Other researchers found that “in almost all cases this decision cannot be made without further context”, because whether a function is vulnerable often depends on how it’s called [Risse2025]. Another study argues that beliefs that LLMs are unreliable at vulnerability detection are “artifacts of context-deprived evaluations” [Li2025].
Whole-project analysis has its own problems. A 2026 study of LLM-based and traditional tools on 24 active open source projects found that “many tools exhibit high false discovery rates on real-world projects/modules”, mostly due to “Truncated interprocedural context and incorrect program-point classification” [Li2026].
Where you can, try a tool on code you know well (including known past vulnerabilities) and compare the validated results.
You can use any existing AI system that can work with code to do this. For example, you might consider directly using an agent (like Goose) to interact with your AI models. However, let’s focus on tools and services specifically for finding and fixing vulnerabilities using AI.
Many organizations with more traditional tools have added modern AI to their products. By “traditional tools” we include Static Application Security Testing (SAST) tools, Dynamic Application Security Testing (DAST) and/or fuzzing tools, and Software Composition Analysis (SCA) tools. Examples of such organizations include Black Duck, Checkmarx, OpenText (Fortify), Snyk, Sonar (SonarQube), Sonatype, and VeraCode.
Closed-source commercial services that focus on using AI to find and fix vulnerabilities, at the time of this writing, include Almanax, Corgea, Cycode, and ZeroPath.
Several frontier AI labs also offer specialized vulnerability-finding and fixing services to outside organizations. At the time of this writing, sometimes access is limited to selected customers or approved defenders. For example:
Open source tools that focus on using AI to find and/or fix vulnerabilities, at the time of this writing, include:
Let’s look more closely at two of these, OpenSSF Alpha-Omega’s Scrutineer and OpenSSF OSS-CRS. They illustrate two ends of a range: a relatively simple set of skills, and a framework for running many CRSs at once. They’re open source software projects of the OpenSSF, which produces this material.
Scrutineer is a tool developed by OpenSSF’s Alpha-Omega. It’s a “local tool for scanning open source repositories for security vulnerabilities and managing the disclosure process. You add a repo by URL, scrutineer runs a pipeline of agent skills against it inside a container, and presents the results in a web UI where you can triage findings, identify maintainers, and track disclosures. The agent CLI is pluggable…” [Nesbitt2026-06].
Scrutineer is intended to be relatively easy to start using. Some aspects of it are especially important:
For more information, see [Nesbitt2026-06] or its website at <https://github.com/alpha-omega-security/scrutineer>.
OpenSSF’s OSS-CRS provides a sophisticated set of capabilities to deeply find and fix vulnerabilities using a variety of techniques. OSS-CRS is a framework for running many different CRSs and combining their techniques. It includes infrastructure that CRSs can share and budget-aware resource management. OSS-CRS is especially helpful when you want to spend significant effort finding and fixing vulnerabilities, to squeeze out as many as is practical.
It’s easier to understand OSS-CRS by understanding its history. DARPA’s AI Cyber Challenge (AIxCC) of 2023-2025 “showed that cyber reasoning systems (CRSs) can go beyond vulnerability discovery to autonomously confirm and patch bugs”. As a research competition it was a success, and the competition results were released as open source software (OSS). However, those systems were “largely unusable outside their original teams, each bound to the competition cloud infrastructure that no longer exists” [Chin2026] [Chin2026-slides].
The solution was OSS-CRS. OSS-CRS builds on the previous AIxCC work to provide a framework for running multiple CRSs and combining their results.
An especially powerful ability of OSS-CRS is its “ensemble” feature. The ensemble feature combines “patches from multiple CRS approaches and [uses] a selection process to pick the one most likely to be correct. The research showed this approach consistently matches or outperforms the best single component in improving semantic correctness, which is hard to eliminate at the single-agent level.” Even so, it’s important to have humans review the proposed changes before implementation [Diecks2026].
OSS-CRS also defines a unified interface for CRS development. A CRS using this interface can run across different environments (both local and remote) without modification.
Here’s how OSS-CRS is intended to be used:
With OSS-CRS, users can decide:
Here are a few OSS-CRS key terms and concepts:
| Term | What it is | Purpose |
|---|---|---|
| Harness | A code wrapper that feeds generated inputs into target functions | Lets bug-finding engines safely execute code, measure coverage, and detect crashes |
| (Fuzzing) Seed | An initial, well-formed input payload provided at the start of a test run | Gives fuzzers a starting baseline to reach deep code logic faster instead of generating random bytes from scratch |
| Proof of Vulnerability (PoV) | A specific input payload that reliably triggers a bug or sanitizer crash (note that definition is more specific than some) | Proves a vulnerability exists and serves as a test case to verify that a proposed code patch fixes the bug |
| Fuzzing/ Fuzzer | The automated technique (and engine) that repeatedly feeds generated inputs into a harness to test execution paths and uncover unexpected crashes | Key capabilities: automated input generation & execution; harness-based path testing; uncovers unexpected crashes |
| Triage | The automated process of filtering, deduplicating, and evaluating raw crashes to confirm they are genuine security flaws rather than benign errors | Key capabilities: crash filtering & deduplication; evaluates raw crash logs; confirms genuine security flaws |
| Patch | A targeted source code modification produced by a CRS to eliminate a vulnerability while preserving all existing intended functionality | Key capabilities: targeted source code modifications; vulnerability elimination; preserves intended functionality |
OSS-CRS is more effective if the project being analyzed has a harness using the OSS-Fuzz format. Many different build systems are used to build software (such as Make, CMake, Autoconf, Bazel, and Meson). This lack of commonality can make it challenging to create tools to correctly analyze them. “OSS-CRS mitigates this by building targets through OSS-Fuzz’s official build flows, inheriting the build environment that each project’s maintainers already support” [Chin2026]. OSS-Fuzz can use one of several fuzzing engines to do its tasks, including libFuzzer, AFL++, and Honggfuzz. If a project doesn’t have an OSS-Fuzz harness, consider using AI to help build one. Ensuring OSS-CRS can build and fuzz a program often improves OSS-CRS results.
You can choose to use the many CRSs already available and ported to run on top of OSS-CRS. You can also create your own CRS (see [crs-bug-finding-template] for more).
OSS-CRS is already capable. “Using OSS-CRS, Team Atlanta discovered twenty-five vulnerabilities across sixteen projects spanning a broad range of software including PHP, U-Boot, memcached, and Apache Ignite 3” [Diecks2026].
Here is a set of short videos demonstrating finding bugs using OSS-CRS combined with LibFuzzer:
Here’s a video showing using OSS-CRS to create a proposed patch:
For more information on OSS-CRS, see: https://openssf.org/projects/oss-crs/
Q1. What does OSS-CRS’s “ensemble” feature do, per the material?