Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Open Source Runs the Internet. The Vulnerability Data Is Split Across 40 Databases

Open source vulnerability information is scattered across dozens of separate databases and disclosure feeds, according to OSV. The gaps between them are where real incidents have started.

TechnologyAnalysisRachel NwosuPublished: 27 September 20265 min readSources 6
Open Source Runs the Internet. The Vulnerability Data Is Split Across 40 Databases

OSV, the Open Source Vulnerabilities database run out of Google, doesn't hold advisories itself. It aggregates them. Its front page lists ecosystem counters that read like a census of everything modern software is built from: npm at 228,732 advisories, Chainguard at 1,023,658, Wolfi at 288,261, MinimOS at 144,169, GIT at 108,035, Debian at 68,356, Ubuntu at 65,370. Those are advisory counts per ecosystem, not unique bugs. The same flaw can appear in several of them.

That overlap is the point. The site calls itself "an open, precise, and distributed approach to producing and consuming vulnerability information for open source", built on the OpenSSF OSV schema. The schema exists so a vulnerability can be mapped "precisely to open source package versions or commit hashes" rather than to vague product names. CISA's entry for OSV is blunt about the tooling cost: it "requires a Google Cloud Platform and Google Group account."

Real incidents show why version-level precision matters.

Log4Shell, CVE-2021-44228, scored 10.0 on CVSS. Ransomware crews and state-sponsored actors kept exploiting it for more than a year after a patch existed, according to a Safeguard analysis of open source risks. The XZ Utils backdoor, CVE-2024-3094, also rated 10.0, was found by chance on 29 March 2024. A contributor using the name "Jia Tan" had spent over two years building commit history and trust, and the code had already reached Debian testing and Fedora 40/41 pre-releases. Heartbleed, CVE-2014-0160, exposed private keys on roughly 500,000 servers in April 2014. The 2017 Equifax breach traced to an unpatched Apache Struts flaw, CVE-2017-5638, and cost $1.4 billion in settlements and remediation.

Patching is not the bottleneck

Safeguard's write-up makes a distinction that gets lost in severity dashboards: the problem is rarely that no patch exists. It is that organisations cannot tell which running services actually call the vulnerable function, so patching gets deprioritised against a backlog where everything looks equally urgent. CISA's Known Exploited Vulnerabilities catalogue, limited to CVEs with confirmed active exploitation, passed 1,300 entries as of 2025. A large share are widely embedded open source components such as Apache, OpenSSL and Spring.

Dependency depth makes that worse. Synk's 2020 State of Open Source Security report, cited by Safeguard, found the average JavaScript project pulls in 683 dependencies, 79% of them transitive. Synopsys's 2024 Open Source Security and Risk Analysis report put open source at 70% to 90% of the code in modern applications. SentinelOne's guide frames the same problem as a governance one: effective programmes need dependency mapping, automated policy enforcement that flags or blocks builds, and risk scoring to decide what gets fixed first when staff and time are short.

Then there is the human layer. In March 2016 the developer Azer Koçulu unpublished left-pad, an 11-line npm package, during a naming dispute. Builds broke across the JavaScript ecosystem within hours until npm reinstated it. Log4j sat inside an estimated hundreds of thousands of applications, yet a small volunteer team maintained it. Lead maintainer Volkan Yazıcı has said publicly that the team handled the Log4Shell response essentially unpaid. There is no SLA on volunteer labour.

Licensing carries its own exposure. The Software Freedom Conservancy sued Vizio in 2021 over GPL source code in its smart TVs. In February 2024 a California appellate court ruled that GPL terms are enforceable as third-party beneficiary contract rights.

Fixing the code is one job. Finding it is another

Tooling around the database layer is consolidating. OSV-Scanner, installed via Go, scans SBOMs, lockfiles, project directories and container images. It ships reusable GitHub workflows so CI/CD pipelines can check newly added dependencies in pull requests and run regular scans across a project. The API answers queries by commit hash or by package version, with example calls for a commit SHA or for jinja2 2.4.1 on PyPI. OSV-Scanner also offers a fix mode with in-place and relock strategies for package-lock.json.

Not every project is chasing the same metric. typed-lm, a Rust project published on GitHub, takes a different route: it turns dense decoder models including Llama, Qwen2, Qwen3, Mistral, Gemma, Gemma2 and Gemma3 into a typed semantic-routing API, returning booleans, choices and scores instead of generated text. Its README claims a single forward pass means milliseconds rather than seconds, and publishes latency tables for a single RTX 3070 with F16 weights. On CPU, it reports a 1024-token prefill falling from 28.23 seconds to 13.13 seconds with CPU flash and MKL enabled, and recommends a GGUF Q4_K_M checkpoint with the mkl feature. Those are project-published figures, not independently verified ones.

The economics underneath all of this are unsettled. Writing in August 2026, developer Debamitro recounted interviewing Christian Hammond, founder and CEO of ReviewBoard, and came away with a counterintuitive finding: companies pay for ReviewBoard "not because it is open source, but despite it being open source." Customers pay for support, and some pay for the hosted version. Hammond also told him that ReviewBoard usage is declining at some companies that are doing away with code reviews, and argued that programming languages, and foundational software generally, should be open source.

That leaves the vulnerability problem where it started: spread across more than forty ecosystem databases, most of them maintained by the same unpaid or thinly funded people shipping the code in the first place. Aggregation helps. It does not close the gap between a published advisory and an organisation that knows whether the affected function is running in production.

Comments 0

Sources

6
  1. 01OSV - Open Source VulnerabilitiesEN
  2. 02Open Source Vulnerabilities (OSV) - CISAEN
  3. 03Open Source Vulnerability Management: A Comprehensive Guide - SentinelOneEN
  4. 04Open Source and Making Money in 2026EN
  5. 05Typed-lm: a Rust jev open source alternativeEN
  6. 065 Risks of Open Source Software (With Real Incidents)EN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.