Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Open Weights Roundup: Nvidia Kumo Tabular Leads a Week of Small, Open Releases

Nvidia released Kumo Tabular on 30 September, an open foundation model for tabular prediction that it says needs no training, tuning or feature engineering and ranks first on four benchmarks. It lands in a week crowded with open-weight releases, from a 164 MB speech model to a 145M-parameter language model pretrained by volunteers.

AI & modelsAnalysisGrace OkonkwoPublished: 30 September 20265 min readSources 9
Open Weights Roundup: Nvidia Kumo Tabular Leads a Week of Small, Open Releases

Nvidia's Kumo Tabular is the biggest name in the batch, but it is not a language model. Given a table with labelled rows, it predicts the labels of new rows in a single forward pass. The company says it comes in three sizes, from 28M to 215M parameters, and was pretrained only on artificial data. It is released under the OpenMDW-1.1 license for commercial use, and Nvidia claims first place on TabArena, BeyondArena, TALENT and ScoringBench.

Small models, small downloads

The most striking artifact of the week is not the most capable one. Fermion Research released Phonon-2 on 30 September, an open English speech recognition model that fits in a 164 MB download. The company says it averages 5.21 percent word error across the seven English sets on the Open ASR Leaderboard, and that every open model scoring better is at least 5.8 times its size. Its weights are CC-BY-4.0, derived from Nvidia's Parakeet TDT 0.6B v3.

The engineering claims are specific: each encoder weight is stored as one of five learned levels, about 2.1 bits per weight, packed as base-3 digits five to a byte. On a MacBook Air, Fermion says, an hour of audio becomes text in about 20 seconds. At the low end of the open-weight spectrum, that size-to-accuracy trade is the whole pitch.

Meanwhile, a volunteer project called Coop published a 145M-parameter decoder-only transformer pretrained without servers or funding, according to its GitHub repository. The setup is unusual: workers download a checkpoint from Hugging Face, run local AdamW steps on a personal data shard, and submit a pseudo-gradient as a pull request against a public Hugging Face dataset repo. A GitHub Actions cron job aggregates submissions every five minutes, though the README notes GitHub's shared scheduler fires anywhere from minutes to a few hours apart.

Coop's earlier stage, a 15M-parameter model trained on TinyStories by volunteers in six days, dropped validation loss from 9.01 to 2.8. The project's own documentation is careful about what that proves: a single validation loss says nothing, so the leaderboard fits a slope over the series, against tokens rather than outer steps, and reports it with a standard error.

Tooling and chips

The infrastructure layer moved too. The UK AI Security Institute and Meridian Labs published Inspect, an open-source framework for frontier AI evaluations, on 30 September. It ships over 200 pre-built evaluations, a web-based Inspect View tool, a VS Code extension, and a sandboxing system that runs untrusted model code in Docker, Kubernetes, Modal, Proxmox or Vagrant. It supports more than 20 model providers, plus local inference with HuggingFace, vLLM and SGLang.

The timing matters because evaluation is now a political subject. On 30 September, The Guardian reported that the US Federal Trade Commission has opened an industry-wide investigation into Anthropic, OpenAI and other AI labs, its first official enforcement action touching rogue AI agents. The FTC plans to compel testimony from executives at Anthropic, OpenAI and the research group Metr, according to multiple reports cited by the paper. All three declined to comment to The Guardian. The New York Post first reported the news.

Then there is the chip layer. Deepseek is releasing open-source programming tools for Huawei's Ascend chips, according to a post on the company's official WeChat channel, as reported by The Decoder on 30 September. The centerpiece is TileLang, a programming language originally developed by researchers at Peking University, which Deepseek argues offers a simpler model than Nvidia's CUDA. Reuters reported that the software includes computation and data-movement libraries, and that Deepseek is making all of it open source. Deepseek says Huawei fully supported the work, and the two optimized a supernode of 128 Ascend 950 chips.

Not everyone buys the framing. Research firm SemiAnalysis has tested OpenAI's Jalapeno inference chip and called the CUDA moat "potentially dead," though it cautions that it only tested scenarios with about 8,000 input tokens and 1,000 output tokens, and has not yet run its AgentX benchmark. In August, the same firm found Nvidia well ahead on agent workloads.

What the releases do not settle

Open weights do not mean open process. Coop's aggregator drops over-stale submissions and clips the rest with a trimmed mean or geometric median before a single Nesterov step. Inspect's sandboxing exists precisely because model code cannot be trusted. Phonon-2's licence inherits from another company's model.

And agreement is not accuracy. Jevstiller, an open-source tool covered by The Register on 29 September, distills Jev outputs into a local model and reports 98 percent agreement with the upstream service while answering familiar requests in as little as 15 ms. Its developers are explicit in their own documentation: "Agreement is not accuracy." That sentence could sit under most of this week's releases.

The political backdrop is less settled still. OpenAI chief executive Sam Altman told reporters at the company's developer day on Tuesday that it will not go public until it can "make confident safety decisions," according to Ars Technica. The company is in talks to raise $30 billion or more at a valuation of about $1.4 trillion, Ars Technica reported, citing people familiar with the matter, and the $30 billion target was first reported by Bloomberg. BBC News reported the same event, noting Altman referred to new agents called dots as "remarkably capable, always-on agents that can handle really anything you can think of."

There is a version of this week in which open weights are simply cheaper distribution. Nvidia wants Kumo Tabular inside enterprise pipelines. Fermion wants Phonon-2 on laptops. Deepseek wants Huawei silicon to have software. Coop wants to show the training loop can run on donated hardware. The common thread is not ideology but arithmetic: smaller artifacts, fewer tokens spent, less infrastructure to rent.

Whether that arithmetic improves safety is a separate question, and this week offered no clean answer.

Comments 0

Sources

9
  1. 01Nvidia Kumo Tabular: Open Foundation Model for Tabular PredictionEN
  2. 02Phonon-2: most accurate open speech recognition model in a 164 MB downloadEN
  3. 03Coop: A small language model pretrained by volunteersEN
  4. 04Inspect: An open-source framework for large language model evaluationsEN
  5. 05US trade regulator opens investigation into AI giantsEN
  6. 06China's AI industry closes ranks as Deepseek ships open-source software for Huawei's Ascend chipsEN
  7. 07Jevstiller: Open-source tool distills Jev so you can run it locallyEN
  8. 08OpenAI delays IPO over AI safety concernsEN
  9. 09OpenAI unveils AI assistant 'dots' while safety worries delay new modelEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.