OpenSSF Finding and Fixing Vulnerabilities Using AI

2.4 Cyber Reasoning System (CRS) history

Any AI system that can work with code can be used to try to find and fix vulnerabilities. Some, however, are more autonomous than others. A cyber reasoning system (CRS) is “a software system that can both detect and repair software vulnerabilities autonomously in a given system under test (SUT)” [Wolff2026]. Here we’ll briefly go through the history of crafting CRSs, because that history influences today.

In 1960, “Lick” Licklider asserted that humans and machines would work together, where “computing machines will do the routinizable work” [Licklider1960]. Today’s AI can be used this way; his prediction was simply early.

DARPA’s 2016 Cyber Grand Challenge (CGC) was “the world’s first all-machine cyber hacking tournament”, in which cyber reasoning systems competed to find and patch vulnerabilities in “custom, never-before-analyzed software”. The winner was Mayhem, developed by ForAllSecure [DARPA2016]. That was years before today’s AI models, however.

In 2023, the US DARPA and ARPA-H launched the two-year competition Artificial Intelligence Cyber Challenge (AIxCC) to build autonomous AI systems that find and fix software vulnerabilities (CRSs) in critical open-source infrastructure. At its beginning, there was some skepticism that it could achieve much on complex targets. When the $4 million grand prize was awarded in August 2025, the finalists had shown that CRSs could find and fix real vulnerabilities in real software, using widely different architectural approaches [Zhang2026].

Of course, fully-autonomous AI systems are not the only way to find vulnerabilities. It’s also possible for humans to provide more direction to AI, or for humans and AI to collaborate throughout the process. Even “fully autonomous” CRS systems, in practice, presume initial human direction and human review of the results.

Quiz

Q1. What defines a “Cyber Reasoning System” (CRS), per the material?

  1. A system that can both detect and repair software vulnerabilities autonomously in a system under test
  2. A system that analyzes vulnerabilities already patched by humans to determine if the vulnerabilities have been correctly fixed
  3. A system that analyzes a sequence of logical statements to determine if the stated assertions correctly lead to the claimed conclusions
  4. A system limited to analyzing closed-source binaries only
Show answer Answer: A
Quiz

Q1. What milestone does the material cite regarding AIxCC?

  1. AIxCC concluded that fully-autonomous CRSs never need human review
  2. AIxCC was canceled in 2025 due to lack of results
  3. AIxCC awarded its grand prize in 2025 after successfully demonstrating autonomous vulnerability-finding-and-fixing systems
  4. AIxCC led to a worldwide ban on the use of open-weight models for vulnerability detection due to concerns some attackers might use AI
Show answer Answer: C