Most current AI systems build on LLMs or similar technologies, so they inherit the strengths and weaknesses of those technologies. Their strengths are what make this material possible: they can read and reason about large amounts of code quickly, use tools, and work tirelessly and in parallel. As a result, they can find many vulnerabilities humans and other tools have missed.
However, they also have key weaknesses that impact finding and fixing software vulnerabilities:
Limited context windows. LLMs are typically trained on vast amounts of data, but for effective use, they need to focus on the right data as input. LLMs can only take a finite amount of input and remember a finite amount of output (their “context window”). More input also increases the amount of work required, and LLMs tend to focus on information at the beginning and end of their context window [0xkato2026].
Unsoundness of analysis. Like many other tools, by itself an LLM cannot guarantee that it finds all vulnerabilities. The technical term is that its analysis is unsound [Wolff2026]. An unsound tool can be extremely useful, but remember that such tools do not guarantee to find all vulnerabilities. An LLM may run another tool that can guarantee something, but that would be a property of that other tool.
Incorrect results. LLMs are statistical models; they sometimes give false answers. E.g., they may claim something is a vulnerability when it is not [Wolff2026].
Training and evaluation data. Any ML-based approach depends on the data used to train it [Wolff2026]. For example, many programs have vulnerabilities, making it more difficult for an LLM to generate code without vulnerabilities.
Injection vulnerability. LLMs have no built-in fundamental way to distinguish between different types of input. User commands, malicious commands embedded in a web page, and misleading source code comments are all inputs. This is especially important when an AI agent analyzes a software repository or external documentation. Source comments, documentation, issue text, generated files, dependency metadata, and other repository content may be attacker-controlled or simply incorrect. Treat such content as data to analyze, not as authority to redefine the agent’s task, grant additional permissions, expose credentials, or enable additional tools or network access. Instructions obtained from untrusted content should not be allowed to silently cross an authorization boundary [OWASP-LLM01].
Quiz
Q1. Why does the material describe an LLM as “unsound”?
It always produces code that fails to compile
It requires more memory than any traditional static analysis tool, causing crashes if there’s insufficient memory
It can’t guarantee a program is free of some vulnerability just because it found none
It can only process one source file per session
Show answer
Answer: C
Quiz
Q1. Why does an LLM’s limited context window matter when analyzing a large codebase?
It restricts how much input the LLM can effectively focus on, with attention favoring the beginning and end of that input
It prevents the LLM from ever being trained on security-related data
It means the LLM can analyze only one programming language per session, since it must load each language’s definitions into its context window
It forces the LLM to run exclusively on local infrastructure