“Enforce automated security assessments consistently in your development processes, including using LLM-powered agents to find vulnerabilities before attackers do” [CSA2026]
“All code (human or AI-generated) should pass LLM-driven security review before merge.” [CSA2026]
Be careful when you run AI agents in CI/CD on proposed changes. When a change comes from an untrusted contributor, its title, description, comments, commit messages, and code are all attacker-controlled. An AI agent treats all of that as input, so an attacker can try to use it to give the agent new instructions (prompt injection).
This isn’t hypothetical. In 2025, Aikido Security reported “PromptPwnd”, where GitHub Actions and GitLab CI/CD workflows embedded information from untrusted issue, pull request, or commit into AI agent prompts. They reported that “At least 5 Fortune 500 companies are impacted” [Daelman2025]. In 2026, Aonan Guan and colleagues showed that a malicious pull request title, issue comment, or hidden HTML comment in an issue could make Anthropic’s Claude Code Security Review, Google’s Gemini CLI Action, and GitHub’s Copilot Agent leak the repository’s own API keys and access tokens [Guan2026]. Note that one of those was a security review tool. As Guan explains, the problem is that “these AI agents are given powerful tools (bash execution, git push, API calls) and secrets (API keys, tokens) in the same runtime that processes untrusted user input.”
In the 2025 “s1ngularity” attack, compromised versions of the Nx build package ran any installed Claude, Gemini, or Amazon Q command-line agent on a developer’s system with its permission checks disabled (e.g., --dangerously-skip-permissions) and used it to search the developer’s system for secrets. StepSecurity called this “the first known case where attackers have turned developer AI assistants into tools for supply chain exploitation” [Kurmi2025]. As we noted earlier, it’s important to run agents inside a sandbox.
If you run an AI agent on untrusted proposed changes:
Don’t give that agent secrets or write-capable tokens. By default, GitHub Actions doesn’t expose secrets to pull requests from forks, but some workflows grant that access anyway, e.g., by using the pull_request_target trigger [Guan2026]. Don’t do that for AI agents.
Give a review agent the fewest tools practical [Daelman2025]. A reviewer of a proposed change usually needs to read code and produce a report; it rarely needs a shell, network access, or permission to push, merge, or edit issues. An agent used to find vulnerabilities may need more capabilities, but reviewing a proposed change from an untrusted proposer is different.
Where an agent needs credentials, use short-lived, narrowly scoped ones.
Treat the agent’s output as untrusted data. Don’t let it automatically trigger privileged actions; put a human gate in front of them (see “Establish a human gate and kill switch”).
Separate analysis from action. A low-privilege agent can analyze and report, while a different, trusted process (or a human) should decide what to do with that report.
Quiz
Quiz
Q1. Why is it risky to run an AI review agent that has repository secrets on pull requests from untrusted contributors?
Text an attacker puts in the pull request can instruct the agent to leak secrets or take other actions
Review agents usually approve changes from new contributors without examining the proposed code closely
Secrets given to the agent become part of the model’s training data and can appear in later responses
Untrusted changes force the agent to use a far larger context window, which sharply raises its costs