OpenSSF Finding and Fixing Vulnerabilities Using AI

2.3 Models

Modern AI systems’ capabilities depend on the models they use. Since this is vital, let’s briefly focus on models to better understand the tools we’re using.

2.3.1 External versus local models

AI models, when executed (for “inferencing”), receive data for processing and reply with results. There are two main locations where this AI data processing occurs:

Some powerful models can only be accessed as an external service. Many models have so many parameters and require so much computation that most users and organizations couldn’t run them on their own systems, even if they wanted to. However, if you use an external service, you must trust both the service provider and the wider ecosystem around it. You need to be satisfied that your data won’t be disclosed or exploited, that the service won’t attack you, and that the service will remain available. That last point can fail even when the provider isn’t at fault: in June 2026, a US government directive briefly forced Anthropic to “abruptly disable Fable 5 and Mythos 5 for all our customers” [Anthropic2026-06-12-Export-Control]. The US Department of Commerce lifted those controls at the end of that same month, but access to Mythos 5 was initially restored only for some US organizations [Capoot2026].

Many organizations use external services. Before using any external service, evaluate the external organization and put in place any necessary contractual agreements. In particular, ensure that sensitive data is not sent to them or ensure that their practices for handling that data are appropriate for your circumstances.

Other AI models can be used locally or as part of an organization’s own service. These models are often less capable, but if run locally, they don’t require users to send their private data to an external service and can be less expensive.

Exactly what you can and can’t do with the model depends largely on its license. That includes the ability to run a model locally at all. So, let’s briefly discuss licenses.

2.3.2 Licenses

There are different ways to license a model. The word “license” means “permission”; a license determines how you can (and can’t) use whatever is licensed. In particular, AI models are much easier to run on local or organizational systems if their licenses are more open.

The Generative AI Commons at the LF AI & Data Foundation has designed and developed the Model Openness Framework (MOF). This is “a comprehensive system for evaluating and classifying the completeness and openness of machine learning models” and is available at <https://isitopen.ai/>. Models are released under various licenses, including the Apache 2.0 license and OpenMDW <https://openmdw.ai/>. Terms you’re especially likely to see when discussing types of licenses are:

Some licenses are close to an open weights license yet have some extra restrictions on use.

NVIDIA argues that “open models and open harnesses are essential because they democratize defensive capabilities, increase transparency for defenders, enable cyber defense while protecting data, and complement frontier closed models with customizable, localized controls…” [NVIDIA2026]

Comparing the capabilities of specific models, including those with differing license terms, is outside our scope and constantly changes anyway. However, [NPR2026] reports that “The most advanced open-weight models are less than a year behind the most advanced closed-weight models.”

Finding and fixing vulnerabilities can be done with both closed and open models.

2.3.3 Specialized models

Some models are specialized. You may want to choose them for some tasks, but only do that where it’s sensible to do that.

All models are better at some tasks than others. Models designed to be good at some tasks tend to be better at those tasks. In addition, some models are designed to perform only specific tasks; this focus means they tend to be smaller and faster, in exchange for being good only at those tasks. There are many ways to make a model good at a particular task, e.g., by giving it additional data in that domain or optimizing it for that task during training.

For example, Cisco’s Antares is a family of small language models (SLMs) specifically built to identify known vulnerabilities in an existing codebase. An SLM is simply the application of LLM approaches to a much smaller number of parameters. Cisco reports that these models “outperform many powerful closed- and open-weight models in this critical security task at a fraction of the cost. And they’re compact enough to run locally” [Karbasi2026].

2.3.4 AI guardrails and intentional limitations

Many models and the larger systems for invoking them, especially many closed models, implement built-in safety guardrails and other limitations intended to prevent the AI system from assisting in “dangerous” activities, including cybersecurity uses like creating attacks, even if the human user requests it.

Unfortunately, such systems can be less useful for defense. Hugging Face discovered in 2026 that it was under a powerful AI-driven attack, and when it tried to analyze its logs, it “first used frontier models behind commercial APIs. This did not work: the analysis required submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers’ safety guardrails, which cannot distinguish an incident responder from an attacker. We ran the forensic analysis instead on zai-org/GLM-5.2, an open-weight model, on our own infrastructure. This had a second benefit: no attacker data, and none of the credentials it referenced, left our environment” [HuggingFace2026].

The best way to validate that a defect is a vulnerability is to create an attack and see if it succeeds. However, this is exactly what attackers do, and AI systems often cannot tell the difference. As a result, AI guardrails can sometimes impede defense. Some external providers that implement guardrails can also offer access to models without them, making them more useful for defensive cybersecurity and other tasks. This typically requires special agreements regarding permitted uses and the limitations on who will be granted this access.

Some AI providers prevent direct access to the model for such use cases and instead perform operations, providing only the final results. For example, Anthropic prevents general end users from interacting directly with their best model without restriction; “instead they will work through purpose-built interfaces that run the model in the background and return only a defined output, such as a list of suggested patches, with abuse-prevention checks meant to keep the model within that scope.” [Kovacs2026-08-24] “Claude Security uses Mythos 5 to scan code you own, and returns detailed findings rather than raw outputs without exposing the model itself [so] defenders can access the capabilities… without the model becoming accessible to those who might misuse it.” [Anthropic2026-08-21] Even in these cases, you may need to specifically request and gain access to these facilities.

Before using any specific AI system, ensure that its limitations will not impede your task. You may need to request less-restricted or specially tailored access. Obtaining these permissions takes time and in some cases may not be granted. If you might do this in the future, it’s important to take time now to gain those permissions before you need them.

Quiz

Q1. Per the material, why might an organization choose a local/organizational AI service over an external one?

  1. Local models always outperform external services on every task
  2. Local models never need to be placed in a sandbox
  3. External services require an open source AI license
  4. It avoids sending private data to an external provider
Show answer Answer: D
Quiz

Q1. Why might a defender need an open-weight or less-restricted AI model when analyzing a real attack?

  1. Closed models refuse any request that mentions security, so they can’t be used for defense
  2. Open-weight models never need a sandbox, since they run on the defender’s own computer systems
  3. Providers’ safety guardrails may block requests that contain real attack commands and payloads
  4. Open-weight models are always more capable than closed frontier models at analyzing security logs
Show answer Answer: C