OpenSSF Finding and Fixing Vulnerabilities Using AI

3.7 Examples of relevant tools and services

The number of relevant tools and services has exploded. We can’t list them all. Here are pointers to specific ones you may find especially useful for this purpose, at least to give you an idea of what’s available.

3.7.1 Challenges of finding and evaluating tools

First, we must acknowledge a problem: it can be challenging to find the tools and services to identify vulnerabilities with AI. There are so many blog posts, papers, and products related to AI and vulnerabilities that it can be difficult to find actual products, services, or systems for finding vulnerabilities in software [Bressers2025] [Rogers2025].

It’s also hard to evaluate them. Many published results come from the tool’s own developers, use different codebases and definitions of success, and report findings rather than validated vulnerabilities.

Be skeptical of benchmark numbers, especially for tools that classify isolated functions as vulnerable or not. Many such benchmarks have inaccurate labels: when researchers manually checked samples of functions labeled “vulnerable” in popular automatically labeled datasets, only 25% to 60% were labeled correctly [Ding2025]. Other researchers found that “in almost all cases this decision cannot be made without further context”, because whether a function is vulnerable often depends on how it’s called [Risse2025]. Another study argues that beliefs that LLMs are unreliable at vulnerability detection are “artifacts of context-deprived evaluations” [Li2025].

Whole-project analysis has its own problems. A 2026 study of LLM-based and traditional tools on 24 active open source projects found that “many tools exhibit high false discovery rates on real-world projects/modules”, mostly due to “Truncated interprocedural context and incorrect program-point classification” [Li2026].

Where you can, try a tool on code you know well (including known past vulnerabilities) and compare the validated results.

3.7.2 Sample services and tools

You can use any existing AI system that can work with code to do this. For example, you might consider directly using an agent (like Goose) to interact with your AI models. However, let’s focus on tools and services specifically for finding and fixing vulnerabilities using AI.

Many organizations with more traditional tools have added modern AI to their products. By “traditional tools” we include Static Application Security Testing (SAST) tools, Dynamic Application Security Testing (DAST) and/or fuzzing tools, and Software Composition Analysis (SCA) tools. Examples of such organizations include Black Duck, Checkmarx, OpenText (Fortify), Snyk, Sonar (SonarQube), Sonatype, and VeraCode.

Closed-source commercial services that focus on using AI to find and fix vulnerabilities, at the time of this writing, include Almanax, Corgea, Cycode, and ZeroPath.

Several frontier AI labs also offer specialized vulnerability-finding and fixing services to outside organizations. At the time of this writing, sometimes access is limited to selected customers or approved defenders. For example:

  1. Anthropic offers Claude Security, discussed earlier.
  2. Google’s CodeMender began as an internal Google DeepMind research project. A version using a more capable model “will be exclusively available to a small set of governments and trusted partners” [Gerstenhaber2026].
  3. OpenAI launched its Daybreak initiative in May 2026, combining OpenAI’s models with its Codex Security tool. At launch, access was tightly controlled, and OpenAI asked interested organizations to request a vulnerability scan or contact its sales team [Lakshmanan2026]. Trail of Bits’ Patch the Planet effort, discussed later, was done in partnership with OpenAI’s Daybreak initiative [TrailofBits2026-06].

Open source tools that focus on using AI to find and/or fix vulnerabilities, at the time of this writing, include:

  1. Alibaba open-code-review <https://github.com/alibaba/open-code-review> [Alibaba2026]
  2. Nullpointer. This focuses on AI-powered pentesting <https://nullpointer.studio/>
  3. OpenSSF Alpha-Omega Scrutineer, a set of skills for finding and fixing vulnerabilities, then reporting them to the external project <https://github.com/alpha-omega-security/scrutineer>
  4. OpenSSF OSS-CRS. This is a meta-tool for creating CRSs, and several CRSs build on it, based on extensive work for AIxCC <https://openssf.org/projects/oss-crs/>
  5. Sashiko. This is a patch review system specifically for the Linux kernel <https://sashiko.dev/>
  6. Visa Vulnerability Agentic Harness <https://github.com/visa/visa-vulnerability-agentic-harness>
  7. Knostic OpenAnt. This analyzes code units reachable from external entry points, then tries to exploit candidate vulnerabilities in sandboxed containers, keeping only what survives. Knostic also offers free scans for open source projects <https://github.com/knostic/OpenAnt> [Korda2026]
  8. OpenAI Codex Security CLI and TypeScript SDK. The client is open source (Apache-2.0), but it uses OpenAI’s Codex Security service, so you need an OpenAI account or API key <https://github.com/openai/codex-security>

Let’s look more closely at two of these, OpenSSF Alpha-Omega’s Scrutineer and OpenSSF OSS-CRS. They illustrate two ends of a range: a relatively simple set of skills, and a framework for running many CRSs at once. They’re open source software projects of the OpenSSF, which produces this material.

3.7.3 Scrutineer

Scrutineer is a tool developed by OpenSSF’s Alpha-Omega. It’s a “local tool for scanning open source repositories for security vulnerabilities and managing the disclosure process. You add a repo by URL, scrutineer runs a pipeline of agent skills against it inside a container, and presents the results in a web UI where you can triage findings, identify maintainers, and track disclosures. The agent CLI is pluggable…” [Nesbitt2026-06].

Scrutineer is intended to be relatively easy to start using. Some aspects of it are especially important:

For more information, see [Nesbitt2026-06] or its website at <https://github.com/alpha-omega-security/scrutineer>.

3.7.4 OSS-CRS

OpenSSF’s OSS-CRS provides a sophisticated set of capabilities to deeply find and fix vulnerabilities using a variety of techniques. OSS-CRS is a framework for running many different CRSs and combining their techniques. It includes infrastructure that CRSs can share and budget-aware resource management. OSS-CRS is especially helpful when you want to spend significant effort finding and fixing vulnerabilities, to squeeze out as many as is practical.

3.7.4.1 OSS-CRS introduction

It’s easier to understand OSS-CRS by understanding its history. DARPA’s AI Cyber Challenge (AIxCC) of 2023-2025 “showed that cyber reasoning systems (CRSs) can go beyond vulnerability discovery to autonomously confirm and patch bugs”. As a research competition it was a success, and the competition results were released as open source software (OSS). However, those systems were “largely unusable outside their original teams, each bound to the competition cloud infrastructure that no longer exists” [Chin2026] [Chin2026-slides].

The solution was OSS-CRS. OSS-CRS builds on the previous AIxCC work to provide a framework for running multiple CRSs and combining their results.

An especially powerful ability of OSS-CRS is its “ensemble” feature. The ensemble feature combines “patches from multiple CRS approaches and [uses] a selection process to pick the one most likely to be correct. The research showed this approach consistently matches or outperforms the best single component in improving semantic correctness, which is hard to eliminate at the single-agent level.” Even so, it’s important to have humans review the proposed changes before implementation [Diecks2026].

OSS-CRS also defines a unified interface for CRS development. A CRS using this interface can run across different environments (both local and remote) without modification.

3.7.4.2 OSS-CRS intended use

Here’s how OSS-CRS is intended to be used:

  1. Multiple CRS techniques are run in parallel. This often involves using fuzzing, static analysis, and LLM-based reasoning, probing the target for bugs, and eventually producing proposed patches.
  2. Ensemble selection picks the best patch. Candidate patches from different CRS approaches are cross-validated.
  3. People review findings and proposed patches.
  4. Verified patches reach the project. Only findings that survive ensemble validation and human review should be submitted.

With OSS-CRS, users can decide:

3.7.4.3 OSS-CRS key terms and concepts

Here are a few OSS-CRS key terms and concepts:

Term What it is Purpose
Harness A code wrapper that feeds generated inputs into target functions Lets bug-finding engines safely execute code, measure coverage, and detect crashes
(Fuzzing) Seed An initial, well-formed input payload provided at the start of a test run Gives fuzzers a starting baseline to reach deep code logic faster instead of generating random bytes from scratch
Proof of Vulnerability (PoV) A specific input payload that reliably triggers a bug or sanitizer crash (note that definition is more specific than some) Proves a vulnerability exists and serves as a test case to verify that a proposed code patch fixes the bug
Fuzzing/ Fuzzer The automated technique (and engine) that repeatedly feeds generated inputs into a harness to test execution paths and uncover unexpected crashes Key capabilities: automated input generation & execution; harness-based path testing; uncovers unexpected crashes
Triage The automated process of filtering, deduplicating, and evaluating raw crashes to confirm they are genuine security flaws rather than benign errors Key capabilities: crash filtering & deduplication; evaluates raw crash logs; confirms genuine security flaws
Patch A targeted source code modification produced by a CRS to eliminate a vulnerability while preserving all existing intended functionality Key capabilities: targeted source code modifications; vulnerability elimination; preserves intended functionality

3.7.4.4 Using OSS-CRS effectively

OSS-CRS is more effective if the project being analyzed has a harness using the OSS-Fuzz format. Many different build systems are used to build software (such as Make, CMake, Autoconf, Bazel, and Meson). This lack of commonality can make it challenging to create tools to correctly analyze them. “OSS-CRS mitigates this by building targets through OSS-Fuzz’s official build flows, inheriting the build environment that each project’s maintainers already support” [Chin2026]. OSS-Fuzz can use one of several fuzzing engines to do its tasks, including libFuzzer, AFL++, and Honggfuzz. If a project doesn’t have an OSS-Fuzz harness, consider using AI to help build one. Ensuring OSS-CRS can build and fuzz a program often improves OSS-CRS results.

You can choose to use the many CRSs already available and ported to run on top of OSS-CRS. You can also create your own CRS (see [crs-bug-finding-template] for more).

OSS-CRS is already capable. “Using OSS-CRS, Team Atlanta discovered twenty-five vulnerabilities across sixteen projects spanning a broad range of software including PHP, U-Boot, memcached, and Apache Ignite 3” [Diecks2026].

3.7.4.5 OSS-CRS demos

Here is a set of short videos demonstrating finding bugs using OSS-CRS combined with LibFuzzer:

  1. 🎬 Set up OSS-CRS
  2. 🎬 Set Variables for CRS Libfuzzer
  3. 🎬 Prepare the CRS Libfuzzer
  4. 🎬 Build a target for the Libfuzzer
  5. 🎬 Run the CRS Libfuzzer
  6. 🎬 Libfuzzer outputs: seeds and Proof of Vulnerabilities

Here’s a video showing using OSS-CRS to create a proposed patch:

🎬 Running a patching CRS

For more information on OSS-CRS, see: https://openssf.org/projects/oss-crs/

Quiz

Q1. What does OSS-CRS’s “ensemble” feature do, per the material?

  1. It combines proposed patches from multiple CRS approaches and picks the best one
  2. It runs a single CRS repeatedly to avoid any conflicting results
  3. It removes the need for any human review of proposed patches
  4. It requires proprietary compute infrastructure available only to DARPA
Show answer Answer: A