Any AI system that can work with code can be used to try to find and fix vulnerabilities. Some, however, are more autonomous than others. A cyber reasoning system (CRS) is “a software system that can both detect and repair software vulnerabilities autonomously in a given system under test (SUT)” [Wolff2026]. Here we’ll briefly go through the history of crafting CRSs, because that history influences today.
In 1960, “Lick” Licklider asserted that humans and machines would work together, where “computing machines will do the routinizable work” [Licklider1960]. Today’s AI can be used this way; his prediction was simply early.
DARPA’s 2016 Cyber Grand Challenge (CGC) was “the world’s first all-machine cyber hacking tournament”, in which cyber reasoning systems competed to find and patch vulnerabilities in “custom, never-before-analyzed software”. The winner was Mayhem, developed by ForAllSecure [DARPA2016]. That was years before today’s AI models, however.
In 2023, the US DARPA and ARPA-H launched the two-year competition Artificial Intelligence Cyber Challenge (AIxCC) to build autonomous AI systems that find and fix software vulnerabilities (CRSs) in critical open-source infrastructure. At its beginning, there was some skepticism that it could achieve much on complex targets. When the $4 million grand prize was awarded in August 2025, the finalists had shown that CRSs could find and fix real vulnerabilities in real software, using widely different architectural approaches [Zhang2026].
Of course, fully-autonomous AI systems are not the only way to find vulnerabilities. It’s also possible for humans to provide more direction to AI, or for humans and AI to collaborate throughout the process. Even “fully autonomous” CRS systems, in practice, presume initial human direction and human review of the results.
Q1. What defines a “Cyber Reasoning System” (CRS), per the material?
Q1. What milestone does the material cite regarding AIxCC?