Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Local 9B model audits 27,489 CVEs, finds 4.2% lack security impact

A new local inference experiment using Cloudflare's 9B Clef Flash model analyzed 27,489 recent vulnerability records, revealing that 4.2% of non-kernel CVEs fail to state a security impact, a gap that undermines automated defense workflows.

AI & modelsAnalysisGrace OkonkwoPublished: 4 October 20268 min readSources 10
Local 9B model audits 27,489 CVEs, finds 4.2% lack security impact

On 4 October 2026, a new study was published on the blog of Jerry Gamblin. It attempts to solve a long-standing classification problem in security operations by running local language models against a massive dataset of recent disclosures.

Security researchers have long struggled to determine whether a published vulnerability record actually explains why a bug matters to a defender. The problem is not just about parsing fields; it is about understanding the prose. Gamblin, who serves on the CVE Program's working groups, emphasizes that the findings are not an official program position. The study uses Ollama's new decision endpoint to run Cloudflare's Clef models on a MacBook, answering two specific questions for every CVE published between August 4 and October 3, 2026: does the description say why the bug matters, and is it clear?

The technical setup relies on a new type of model interface introduced in Ollama 0.35. These models do not generate open-ended text. Instead, they accept a state and a set of typed questions, returning choices, yes-or-no probabilities, and scores.

The results are striking in their simplicity. Outside of the Linux kernel, 989 of 23,736 descriptions, or 4.2 percent, never state what an attacker could do. That is roughly one in every 24 records. The Linux kernel presents a different case due to its published policy, which leaves impact assessment to the user. Consequently, 3,661 of its 3,753 records do not state an impact. When combined, the all-CVE figure rises to 4,650 out of 27,489, representing 16.9 percent of the dataset.

Two-step verification reduces false positives

Running a single 9B model is not enough to draw definitive conclusions. The author acknowledges that Clef Flash on its own is too strict. To mitigate this, the study employed a second model, the full 27B Clef model, to re-read everything flagged as having "no impact" outside the kernel. This two-step process is designed to filter out false negatives where a smaller model might miss a subtle implication of risk.

The 27B model overturned 1,198 of the 2,187 initial flags, a rate of 54.8 percent. This significant correction rate highlights the difficulty of the task. To validate the final output, the author compared the model's answers against a set of 129 published CVEs that he labeled by hand. The two-step answer caught 15 of the 16 records he identified as stating no impact. Crucially, it never called one of his "yes" labels a "no," indicating a high precision in the negative detection class.

Clef Flash answered both questions for every new CVE at a median of under a second each on a laptop. There was no API key required and no bill generated. This speed and cost profile make it feasible for organizations to audit their own vulnerability management pipelines without relying on cloud-based services or proprietary APIs.

The context of local inference

This study fits into a broader trend of bringing advanced AI capabilities to local hardware. The use of Ollama's decision endpoint is a specific example of this shift. The endpoint allows developers to treat language models as classifiers rather than text generators. This is particularly useful for tasks that require structured outputs, such as security triage, where the cost of a hallucinated sentence is far higher than the cost of a misclassified label.

The models themselves are fine-tunes of open-weights architectures. Clef is a 27B model fine-tuned from Qwen 3.8, while Clef Flash is a 9B model fine-tuned from Qwen 3.5. Both are hosted on Cloudflare's Workers AI platform, but the study was conducted entirely on a local laptop. This distinction is important because it demonstrates that the utility of these models does not depend on centralized infrastructure. A security team can pull the model and run the audit on their own hardware, ensuring that sensitive vulnerability data never leaves their network.

The dataset covers a 60-day window, from August 4 to October 3, 2026. This recency is critical for understanding the current state of vulnerability disclosure practices. The 27,489 records represent a significant volume of data, but the local inference approach makes it manageable. The median response time of under a second per record means the entire audit could be completed in a matter of hours on standard consumer hardware.

Implications for defense automation

The finding that 4.2 percent of non-kernel CVEs lack a stated security impact has direct implications for automated defense systems. Many security tools rely on CVE descriptions to prioritize patching efforts. If a description does not explain the impact, the tool cannot accurately assess the risk. This gap forces defenders to rely on manual review or external databases, which can introduce delays and inconsistencies.

The study suggests that local models can help bridge this gap by providing a consistent, scalable method for assessing the quality of CVE records. The two-step verification process, using a smaller model for initial filtering and a larger model for re-evaluation, offers a balance between speed and accuracy. This approach could be adapted to other domains where structured classification is needed, such as log analysis or incident response triage.

The author's background in the CVE Program's working groups adds weight to the findings, even though he explicitly states that the results are not an official position. This distinction is important because it allows for independent scrutiny and criticism. The transparency of the methodology, including the use of open-weights models and the publication of the request structure, enables other researchers to replicate the study and verify the results.

Comparing model capabilities

The choice of Clef Flash and Clef for this task is not arbitrary. The decision endpoint in Ollama 0.35 was designed specifically for this type of classification work. The models are fine-tuned to handle typed questions, which allows for precise control over the output. This is in contrast to general-purpose chat models, which may produce verbose or inconsistent responses that are difficult to parse automatically.

The 9B size of Clef Flash makes it suitable for local execution on consumer hardware. The 27B size of the full Clef model requires more resources but provides higher accuracy for the re-evaluation step. This tiered approach allows organizations to scale their audits based on their hardware capabilities and accuracy requirements. A small team with limited resources could use only the 9B model for a quick scan, while a larger organization could use the two-step process for a more rigorous audit.

The study also highlights the importance of data quality in security operations. The 4.2 percent gap in impact statements is not just a statistical curiosity; it is a practical obstacle to effective defense. By identifying this gap, the study provides a baseline for measuring improvements in CVE description quality. It also suggests a tool for monitoring this quality over time, allowing organizations to track whether disclosure practices are improving or deteriorating.

Reproducibility and open source

A key strength of this study is its reproducibility. The models are available on Hugging Face, the code for the Ollama decision endpoint is open source, and the dataset of CVEs is publicly available. This transparency allows other researchers to replicate the study, verify the results, and extend the analysis to other time periods or model families. The use of open-weights models also ensures that the methodology is not dependent on a single vendor's proprietary API.

The study's focus on local inference is particularly relevant in the current climate of data privacy and security. Many organizations are hesitant to send sensitive data to cloud-based AI services due to concerns about data leakage or vendor lock-in. The ability to run these models locally addresses these concerns, providing a secure and private method for auditing vulnerability records. This is a significant advantage for organizations in regulated industries, such as finance or healthcare, where data protection is paramount.

The findings have broader implications for the field of AI-driven security. As language models become more capable, their role in security operations is likely to expand. The ability to use local models for classification tasks like this one suggests that AI can be a powerful tool for improving the quality and efficiency of security processes. However, it also highlights the need for careful validation and verification, as demonstrated by the two-step process used in this study.

Future directions

The study opens up several avenues for future research. One direction is to extend the analysis to other types of security data, such as incident reports or threat intelligence feeds. The same classification approach could be used to assess the quality and clarity of these documents, providing a consistent method for prioritizing response efforts. Another direction is to investigate the use of different model architectures or fine-tuning techniques to improve the accuracy of the classification task.

The study also raises questions about the role of AI in the CVE program itself. While the author states that the findings are not an official position, the results could inform future policies and guidelines for CVE description quality. The use of AI to audit these descriptions could become a standard part of the CVE process, ensuring that all records meet a minimum threshold of clarity and impact statement. This would benefit defenders by providing more reliable data for their decision-making processes.

This study demonstrates the potential of local language models for security analysis. By using a two-step verification process with open-weights models, the author provides a solid and reproducible method for assessing the quality of CVE descriptions. The findings highlight a significant gap in current disclosure practices and offer a tool for addressing this gap. As AI continues to evolve, its role in security operations is likely to grow, and studies like this one will be essential for understanding and managing that evolution.

Comments 0

Sources

10
  1. 01Local models on 27,489 CVEs: 4.2% omit security impact (ex-kernel)EN
  2. 02Google's new Gemini tiers cut free users to its weakest modelEN
  3. 03Gemini app limiting what models free and AI Plus users can accessEN
  4. 04OpenAI's latest features take direct aim at the app store modelEN
  5. 05Anthropic asks Claude users to share voice data for AI model trainingEN
  6. 06Ideogram 4.5: The most precise edit modelEN
  7. 07What happens when an AI model is put in a "pain" state?EN
  8. 08Benchmarking CV and depth-estimation algorithmsEN
  9. 09Language Model "Shape" – Alex L. ZhangEN
  10. 10No Model Required: Text Entropy Rate Filtering Mitigates Iterative Fine-TuningEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.