OpenSSF Finding and Fixing Vulnerabilities Using AI

4.1 Preparing AI

A key step is “preparing AI”, that is, preparing the AI system itself. We list this first, because you’ll probably want to use AI to help you perform the other preparation steps. Preparing AI means selecting the AI tools, including possibly their models, harnesses, and so on. There’s no requirement that you choose a single system to do this. In fact, different systems have different strengths and costs. What’s more, there’s no general guarantee that an AI system will find all vulnerabilities, so using multiple systems over the long term has its advantages.

Different AI systems vary in cost as measured in both money and time. For example, older models and smaller models (including SLMs) are often less capable but can also be less expensive. In practice, it may be helpful to choose less expensive and faster AI systems first to implement vulnerability finding and fixing, focusing on vulnerabilities that are simpler to find and fix. Taking this approach means that less expensive models quickly handle the vulnerabilities that are easy to find, fix, and verify. Once those vulnerabilities are addressed, more capable AI systems can focus on the vulnerabilities best addressed by them.

That said, don’t use especially weak models or models that are a poor match. You may want to set the models’ reasoning (effort) level to relatively high levels. Many of these tasks are challenging for AI (and for humans).

In many AI systems, output tokens cost significantly more than input tokens. Taking steps to eliminate unnecessary output can reduce costs and time. You can sometimes do this by asking for concise results, using strict formatting requirements such as JSON schemas or bullet points, or by defining a hard ceiling on API calls. However, the challenge is to avoid unnecessary output; you’ll need enough output to have results and verify the work.

As noted earlier, a key decision is whether to use an external (remote) system, since that means the external system will receive the data for processing. As [Kholoosi2025] notes, “due to internal policies, LLMs hosted on [external servers sometimes] cannot be used”.

You’ll generally want the AI system to be able to record memories of preferences, system information, and so on. For a remote system, that typically means using an account rather than anonymous access. Recorded memories let the AI build on earlier work, and can improve its results, but they need care.

Persistent AI memory should be treated as another input, not as authoritative truth. Stored information may become stale, preserve an incorrect earlier conclusion, contain sensitive information, or be influenced by untrusted inputs. When memory affects a security decision, prefer current source code, configuration, threat-model information, and reproducible evidence over remembered conclusions, and retain provenance for stored information where practical [OWASP-ASI06].

We’ll need to control the AI, including putting it in a sandbox. However, how to do that well depends on the system under review, so we’ll discuss that in more detail once we discuss how to determine the system’s security requirements via a threat model.

Quiz

Q1. Why might it help to use less expensive, faster AI systems first, according to the material?

  1. They eliminate the value of using more-expensive AI systems
  2. They eliminate the need for any sandbox environment entirely
  3. They automatically outperform frontier models on every possible task
  4. They quickly resolve simple bugs, freeing costlier systems for harder ones
Show answer Answer: D
Quiz

Q1. How does the material say an AI’s persistent memory should be treated?

  1. As authoritative truth that overrides current source code and configuration
  2. As something that should never be enabled under any circumstances
  3. As another input that may be stale or shaped by untrusted data
  4. As a full, reliable substitute for the entire CI/CD pipeline
Show answer Answer: C