Before using an AI system to make changes to code, ensure you have a robust Continuous Integration / Continuous Delivery or Deployment (CI/CD) system that can rigorously check its work.
An AI system generally works best “when it’s able to check its own work with another tool. We refer to this class of tool as a ‘task verifier’: a trusted method of confirming whether an AI agent’s output actually achieves its goal. Task verifiers give the agent real-time feedback… allowing it to iterate deeply until it succeeds” [Anthropic2026-03].
Because task verifiers are trusted to accept or reject an agent’s work, in a sense they become part of the trusted computing base for the workflow. A broken test, an incomplete harness, a compromised build environment, or a verifier that checks the wrong property can confidently approve an incorrect result. Where practical, keep verifiers deterministic, make explicit what property each verifier actually proves, and periodically test them using known-good and known-bad cases.
Thankfully, any software that needs to work correctly should already have mechanisms that support such checking. Any such software should have a CI/CD pipeline to build, test, and deliver changed results:
Continuous Integration (CI) means that code is frequently merged and verified before acceptance.
“CD” can mean either Continuous Delivery (where the verified results are automatically prepared for deployment) or Continuous Deployment (where the verified results are automatically released to live environments).
Having a good CI/CD process has always been important, but it’s even more so with AI. A process that relies on people remembering to test, or a manual testing process, is not equipped to handle the large number of vulnerabilities and fixes required by today’s systems.
The CI/CD process needs to be high quality to reduce the likelihood of breaking functionality or introducing vulnerabilities. For example:
Ensure that you have a good automated test suite
Include negative tests (these are tests to verify that what should not happen doesn’t happen)
Have good statement coverage (e.g., 90%-100%) [0xkato2024] and branch coverage
Don’t just output “test failed”; report which test(s) failed, the expected results, and the actual results. This is necessary to speed response
Use linters to detect possible defects (some of which may be vulnerabilities)
Include tools to detect likely vulnerabilities, such as static application security testing (SAST) tools
Take steps to minimize false positives, and especially work to counter repeat false positives. E.g., mark confirmed false positives in the source code, such as with suppression comments, so tools don’t report them again
Speed matters in these processes. It doesn’t matter if a vulnerability is known; what matters is deploying the fix before an attacker exploits it:
Make CI time acceptable (including testing). “If regression testing takes a day, you cannot get to a two-hour SLA without skipping it, and the bugs you ship when you skip regression testing tend to be worse than the bugs you were trying to patch.” [Bourzikas2026] Thankfully, most of these tasks are easily parallelizable. Break them down and run them in parallel to reduce wall-clock time to acceptable times. You should be thinking about total minutes, at worst an hour, not a day or longer for most projects. Obviously, such short times can be challenging for operational technology (OT) or software that controls physical devices, but the principle of reducing CI time still applies. If CI takes too long, CI becomes the bottleneck.
Speed deployment. “Defenders need a continuous operating pipeline that moves from signal to context to action with minimal delay.” [CrowdStrike2026-FiveSteps]
If your CI/CD process may receive information from untrusted users, you need to protect against that. That definitely includes the case where an AI agent is reviewing proposals as part of CI/CD. We’ll discuss that later in the section evaluate merge/pull requests.
Quiz
Q1. What is a “task verifier” as defined in the material?
A human who manually re-reads every line of AI-generated code
A tool that checks only for open source license violations
A separate AI model used solely to write documentation
A trusted method for confirming whether an agent’s output achieves its goal
Show answer
Answer: D
Quiz
Q1. Why does the material say that a broken test or a compromised build environment is especially problematic in AI-driven workflows?
It has little effect, since the AI double-checks its own work
A broken task verifier may repeatedly approve bad results
It affects only performance and never has any effect on correctness
CI/CD pipelines have nothing to do with AI-assisted vulnerability work