OpenSSF Finding and Fixing Vulnerabilities Using AI

4.6 Preparing CI/CD

Before using an AI system to make changes to code, ensure you have a robust Continuous Integration / Continuous Delivery or Deployment (CI/CD) system that can rigorously check its work.

An AI system generally works best “when it’s able to check its own work with another tool. We refer to this class of tool as a ‘task verifier’: a trusted method of confirming whether an AI agent’s output actually achieves its goal. Task verifiers give the agent real-time feedback… allowing it to iterate deeply until it succeeds” [Anthropic2026-03].

Because task verifiers are trusted to accept or reject an agent’s work, in a sense they become part of the trusted computing base for the workflow. A broken test, an incomplete harness, a compromised build environment, or a verifier that checks the wrong property can confidently approve an incorrect result. Where practical, keep verifiers deterministic, make explicit what property each verifier actually proves, and periodically test them using known-good and known-bad cases.

Thankfully, any software that needs to work correctly should already have mechanisms that support such checking. Any such software should have a CI/CD pipeline to build, test, and deliver changed results:

Having a good CI/CD process has always been important, but it’s even more so with AI. A process that relies on people remembering to test, or a manual testing process, is not equipped to handle the large number of vulnerabilities and fixes required by today’s systems.

The CI/CD process needs to be high quality to reduce the likelihood of breaking functionality or introducing vulnerabilities. For example:

Speed matters in these processes. It doesn’t matter if a vulnerability is known; what matters is deploying the fix before an attacker exploits it:

If your CI/CD process may receive information from untrusted users, you need to protect against that. That definitely includes the case where an AI agent is reviewing proposals as part of CI/CD. We’ll discuss that later in the section evaluate merge/pull requests.

Quiz

Q1. What is a “task verifier” as defined in the material?

  1. A human who manually re-reads every line of AI-generated code
  2. A tool that checks only for open source license violations
  3. A separate AI model used solely to write documentation
  4. A trusted method for confirming whether an agent’s output achieves its goal
Show answer Answer: D
Quiz

Q1. Why does the material say that a broken test or a compromised build environment is especially problematic in AI-driven workflows?

  1. It has little effect, since the AI double-checks its own work
  2. A broken task verifier may repeatedly approve bad results
  3. It affects only performance and never has any effect on correctness
  4. CI/CD pipelines have nothing to do with AI-assisted vulnerability work
Show answer Answer: B