3.5 Why AI tends to be more effective if guided by processes
You can find and fix some vulnerabilities with simple prompts, as noted above. In particular, more powerful AI models can sometimes identify and fix vulnerabilities without additional help or specialized processes. However, if your goal is thoroughness, many analysts report that guiding an AI makes the AI much more effective at finding and fixing vulnerabilities.
For example, [Wolff2026] states that “simply asking a generic coding agent to ‘find bugs’ in a large repository results in model drift, context window compaction, and high false-positive rates… The highest-performing defensive systems combine raw AI reasoning with rigid engineering frameworks (harnesses) that handle file selection, environment setup, and tool execution.” An AI that tries to “read all the code” at once has the same problem a human might have; it can become overwhelmed with the data it’s being asked to peruse. AI systems’ limited context windows can make them less effective when used this way.
Similarly, [Grinstead2026-05] noted that “you can start with very simple prompting…” but also that “through iteration we’ve built out a lot of orchestration and tooling to optimize and scale the pipeline.” [Carlini2026] found that asking “each agent to focus on a different file in the project” was more effective, since this “reduces the likelihood that we will find the same bug hundreds of times”. Instead of processing every file for each software project, they first asked Claude to rank how likely each file in the project was to contain interesting bugs, then prioritized the files most likely to contain them.
Derek Zimmer [Zimmer2026] believes that one reason for this is that it’s easy for an AI to “run things out-of-order”. AI may try to anticipate the next step, and what it guesses may be wrong. Having an AI agent perform a specific task, using previously created data, enables it to focus its attention and effort on that task. In particular, adding specialized processes can make less-powerful AI models (including less-expensive ones) far more effective.
[Bourzikas2026] reported 4 lessons, each pointing to the value of a harness to manage the overall process:
“Narrow scope produces better findings. Telling the model “Find vulnerabilities in this repository” makes it wander. Telling it “Look for command injection in this specific function, with this trust boundary above it, here’s the architecture document and here’s prior coverage of this area” makes it do something much closer to what a researcher would actually do.
Adversarial review reduces noise. Adding a second agent between the initial finding and the queue - one with a different prompt, a different model, and no ability to generate its own findings - catches a lot of the noise that the first agent would miss if it just checked its own work. It turns out that putting two agents in deliberate disagreement is way more effective than just telling one agent to be careful.
Splitting the chain across agents produces better reasoning. Asking “Is this code buggy?” and “Can an attacker actually reach this bug from outside the system?” are two different questions, and the model is better at each one when you ask them separately, because each question is narrower than the combined version.
Parallel, narrow tasks beat a single exhaustive agent. Coverage improves when many agents work on tightly scoped questions and we deduplicate the results afterward, rather than asking one agent to be exhaustive.”
At the time of writing, exactly how to best guide AI is under evaluation. Different groups use different approaches, and it’s likely that some approaches are better suited to certain types of vulnerabilities. We’ll further discuss approaches later, but as an example, here’s the approach described by [Bourzikas2026]:
Recon: An agent reads the repository from the top down, fans out to subagents responsible for each subsystem, and produces an architecture document covering build commands, trust boundaries, entry points, and likely attack surface. It also generates the initial task queue for the next stage. Gives every downstream agent shared context. Cuts the wander problem.
Hunt: Each task is one attack class paired with a scope hint. Hunters (the agents that actually look for bugs) run concurrently, typically around fifty at once, each fanning out to a handful of exploration subagents. Each hunter has access to tools that compile and run proof-of-concept code in a per-task scratch directory. This is where most of the work happens. Many narrow tasks in parallel, not one exhaustive agent.
Validate: An independent agent re-reads the code and tries to disprove the original finding. It uses a different prompt and cannot emit new findings of its own. Catches a meaningful fraction of the noise the hunter wouldn’t catch when reviewing its own work.
Gapfill: Hunters flag areas they touched but didn’t cover thoroughly. Those areas get re-queued for another pass. Counteracts the model’s tendency to drift toward attack classes it has already had success with.
Dedupe: Findings that share the same root cause collapse into a single record. Variant analysis is a feature, not a way to inflate the queue with duplicates.
Trace: For each confirmed finding in a shared library, a tracer agent fans out (one instance per consumer repository), uses a cross-repo symbol index, and decides whether attacker-controlled input actually reaches the bug from outside the system. Turns “there is a flaw” into “there is a reachable vulnerability.” This is the stage that matters most.
Feedback: Reachable traces become new hunt tasks in the consumer repositories where the bug is actually exposed. Closes the loop. The pipeline gets better as it runs.
Report: An agent writes a structured report against a predefined schema, validates it against that schema, and submits it to an ingest API. Output is queryable data, not free-form prose.
Of course, many others outline some sort of process. Anthropic reports that teams finding and fixing the most vulnerabilities ended up with some variation of the following steps:
“Threat model: Decide what counts as a vulnerability before you start scanning.
Sandbox: Build a sandbox environment to isolate agents and prove exploits.
Discovery: Have models look for vulnerabilities in your source code.
Verification: Independently confirm which findings are actually exploitable.
Triage: De-duplicate findings, assign severity, and prioritize what needs fixing.
Patching: Apply the fix, confirm the vulnerability is nullified, and search for variants.” [Yan2026]
As we’ll further discuss later, a key task is validation. Something may look like a vulnerability but be unexploitable. The best way to validate a vulnerability is to generate an exploit that demonstrates it is indeed a vulnerability. A working exploit demonstrates that existing defenses wouldn’t prevent the attack [Carlini2026]. This also explains why hardening is so important; if a project hardens its software against attack, it can systematically prevent many problems from becoming vulnerabilities. The Linux kernel, for example, has various defense-in-depth measures that have prevented the exploitation of many potential problems identified by even advanced AI models [Carlini2026].
In 2026, AI became far more effective at finding vulnerabilities. This was due to a combination of improved AI models and improved techniques for harnessing them [Grinstead2026-05]. Using either one is better than using none, but many analysts report that it’s best to combine them.
Of course, once findings (potential vulnerabilities) are found, and then validated as vulnerabilities, they need to be fixed.
🎬See the video “Types of CRS” for an illustration of how vulnerability reports generated by one system can flow into another system designed to fix them.
There is a risk that overprescribing an approach to an AI may overconstrain it, causing it to ignore problems it would otherwise find. It’s also possible that future AI models will be so good that aiding them with processes won’t help. However, since many analysts report that aiding AI models does help, we’ll discuss doing that combination in this material.
Quiz
Q1. Per the material, what can happen when you simply ask a generic coding agent to “find bugs” in a large repository by reading all of its code?
It always refuses the task, citing safety guardrails
It can cause model drift, context compaction, and high false-positive rates
It automatically escalates its own tool permissions and access
It produces final results requiring no further human review
Show answer
Answer: B
Quiz
Q1. What lesson did Cloudflare researchers report about narrowing an AI’s task scope?
Broad instructions like “find vulnerabilities in this repository” and encouragement to analyze everything at once tends to work best
Scope has no measurable effect on the quality of findings at all
Only one agent should ever run at a time, to avoid duplicates
Narrow scope for each AI run, such as one function with a trust boundary, produces the best findings