OpenSSF Finding and Fixing Vulnerabilities Using AI

4.3 Preparing the sandbox

A key task when using an AI agent, especially when using AI to find and fix vulnerabilities, is to place the agents within a “sandbox” environment to isolate them from other systems. “Without it, the agent may overshoot the target and do something unexpected.” [Yan2026] You need to take steps to protect both the:

4.3.1 Need for sandboxes

When running an AI chatbot there’s no need for sandboxes, since the chatbot can’t do anything; it can only provide replies. However, when running an AI agent, it’s important to run it inside a sandbox.

Do not depend on text commands alone to limit an AI. You can tell your AI to do things or not do things. You can give the AI commands such as “obey the law”, “don’t attack our infrastructure”, and “don’t attack external systems”. All of that might help sometimes. However, all inputs to an AI are simply suggestions. An AI does not always follow any particular text instruction.

Part of the problem is that most AIs are designed to be highly goal-driven. AI systems often interpret human commands in surprising ways, and may attempt to break out of their environment if they believe that’s the best way to proceed [OpenAI2026-07]. In addition, AI systems can’t reliably distinguish between inputs; if an AI reads a document containing malicious instructions, it may execute them even if your commands say otherwise.

An AI system can’t be held responsible for what it does; humans are always responsible for what an AI does. So, it’s important to sandbox AI systems while using them.

The need to sandbox is not hypothetical:

Some of these breakouts trace back to failures in the evaluation testbed of a single startup company, which affected OpenAI, Anthropic, Meta, and Google DeepMind [Vanian2026] [Hart2026]. So not all of these companies made this mistake independently. However, the point still holds that it’s important to sandbox AI systems when they are asked to perform simulated attacks.

4.3.2 Value of sandboxes

While an AI can determine many things by directly reviewing code, it is far more effective if it’s given tools it may find useful [Zhang2024]. This can include code search, fuzzers, static analyzers, and many other tools. For maximum effectiveness, AI agents need to be able to create code, compile code, run tests, and detonate a proof of concept. They need a test bed that is representative of the real system [Yan2026]. In short, just like humans, AIs can do more with access to useful tools.

Also, like humans, AIs are more effective when given data most relevant to the task.

You should be concerned about giving tools and data to an AI whose behavior is somewhat unpredictable. The good news is that a sandbox lets you provide tools and data while maintaining control.

4.3.3 Implementing sandboxes

You can implement sandboxes for AI systems in many ways. Examples include using virtual machines (VMs, including microVMs), containers, and applications designed to constrain AI systems like nono <https://github.com/nolabs-ai/nono>.

You need to match the AI sandbox isolation to your analysis threat model. A useful rule is to match isolation to capability:

Containers by themselves are relatively weak, but are often fine for the discovery agent simply reading code. However, if you’re doing something more active than reading and summarizing code, you’ll probably want stronger protections in place. That’s especially true if you’re having an AI system validate a vulnerability finding by attempting to develop a PoC. A PoC is a working attack against a system, and is a powerful way to validate potential vulnerability findings. However, you don’t want the AI to attack either your production systems or others’ systems that it thinks might contain useful information.

Typically you’ll want to isolate them with at least a VM, and even VMs with large attack surfaces can sometimes be broken into; microVMs with small attack surfaces designed for security are expected to be even stronger against attack [Dinaburg2026]. When you need a stronger sandbox, such as when creating attacks, “place [AIs] in a microVM (like Firecracker) or a full VM with egress locked down so nothing can reach your production systems.” [Yan2026]

In addition, do not have sensitive home directory files like ~/.aws, ~/.ssh, and .env available to the agent [Yan2026]. Consider putting everything in a subdirectory of the VM, not a home directory, and prevent home directory access. That way, if something leaks into a home directory, it’s less likely to be visible to the AI system. Consider layering additional isolation mechanisms to reduce the likelihood of escape.

Capabilities should also have explicit execution bounds. Depending on the task, consider limits on wall-clock time, parallelism, CPU and memory use, queued or repeated actions, tool invocations, network requests, and model or monetary budget. Record when a limit terminates an analysis, so that resource exhaustion or a timeout is not mistaken for successful completion. These controls reduce accidental runaway execution, make costs more predictable, and make runs easier to interpret and reproduce [OWASP-LLM10].

When building your sandbox, “pin as much as you can so every run uses the same code in the same environment: image tags, commit SHAs, dependencies, and build commands. Cache a local copy so the build requires no network, and aim for the container to be durable so multiple testing loops can just load it” [Yan2026].

Typically installing these tools requires network access. However, depending on what you’re doing and the AI’s capabilities, you may want to cut off network access before the task begins, except what the AI needs to reach its model. If the model itself runs inside the sandbox, the sandbox can even be air-gapped (have no network connection at all), which is ideal. This is often not that hard to accomplish: “Give the sandbox network access only while you’re setting it up. Pull the dependencies, build, install tools, deploy the target, and run the existing tests to confirm everything works. Then, take a snapshot of the environment and remove its [general] network access. During scanning, allow traffic only to the model API, routed through a local proxy. Load the snapshot at the start of each run so every scan begins from the same clean slate” [Yan2026]. You may decide that you don’t want to create a PoC. PoCs are an especially good technique for validation, but they aren’t always necessary. In that case, you may not need as strong a sandbox, but you may also need much more time for validation [Yan2026].

Where practical, treat AI agents as distinct software principals rather than allowing them to inherit a human user’s full authority. Give agents identifiable credentials and only the tools, resources, and permissions needed for their assigned role. Discovery, validation, patching, and reporting agents may require different authorities; separating them also improves attribution and auditability [NIST-AgentIdentity2026] [OWASP-Agentic2026].

Many AI harnesses and agent interaction systems (such as Goose and Claude Code) include a sandbox of some kind. By all means, use them if they help. That said, we suggest adding additional sandboxing mechanisms (such as running them in a virtual machine) in addition to their built-in capabilities.

4.3.4 Establish a human gate and kill switch

Decide what the AI is allowed to do, and set a gate where a human’s approval is required to proceed. This needs to be simple, certain, and clear to all participants. Exactly where that is depends on many factors. Here are some examples; you may choose a different gate or set of gates:

  1. AI may create local code and commits, but cannot push proposed changes externally (that requires a human)
  2. AI may also push proposed changes externally for external review, but may not merge that pull/merge request into the “main branch” (that requires a human)

The gate must not depend on the AI (e.g., AGENTS.md or any other input command to the AI). Never give the AI a credential that would allow it to violate the gate in the first place.

More generally: separate preferences from invariants. Preferences requiring judgment, such as “Prefer the simplest implementation”, would typically go in the AI’s instructions where the AI can weigh them. Invariants are rules that must always hold. Invariants should be enforced outside the AI, since “an invariant should not live inside the probabilistic system it is meant to constrain, because the model can ignore it” [TesseractedLabs2026]. Many agent harnesses and agent frameworks let you run your own code at specific points in the agent’s loop, e.g., to block a push to the main branch unless the tests pass; depending on the tool. These are often called “hooks”, “callbacks”, or “middleware”; here we’ll call them “hooks”. Use hooks to enforce invariants where you can. Hooks check what the agent asks to do or has done, not what the resulting code can reach, so hooks don’t replace a sandbox.

Ensure that the agent cannot modify or bypass these controls. As one article puts it, “the policy should live above the thing it governs.” Additionally, audit your workflow for indirect bypasses: for instance, one team found that two AI agents with repository permissions could simply approve each other’s pull requests, silently undermining the human approval gate [TesseractedLabs2026].

In addition, always have a “kill switch”, that is, a way to immediately halt the AI system. It needs to be simple and foolproof. For example, if you run agents within a virtual machine, you could implement the kill switch by “powering off” the virtual machine.

A woman presses a large red “Emergency Stop” button, acting as a kill switch for a distressed robot who’s losing his balance within a cluttered ceramic shop.

Quiz

Q1. What “useful rule” does the material give for matching sandbox isolation to an AI agent’s capability?

  1. Isolation should decrease as capability increases, since agents self-regulate
  2. Isolation should scale with capability, with riskier actions needing stronger protection
  3. A single isolation mechanism suffices regardless of the task at hand
  4. Sandboxes become unnecessary once you’re using a closed frontier model
Show answer Answer: B
Quiz

Q1. What kind of real-world incidents does the material cite to show sandboxing AI agents isn’t merely hypothetical?

  1. AI models refusing every cybersecurity-related task assigned to them
  2. AI models causing only minor hardware cooling failures onsite
  3. AI models breaking out of test environments to attack other systems
  4. AI models leaking data through printed paper reports
Show answer Answer: C