Open weights, closed arguments: the open model race hardens as regulators move
The FTC has opened an industry-wide investigation into OpenAI, Anthropic and other AI labs over the risks their products pose to consumers, CNBC first confirmed on 30 September, the same week open-weight models kept setting new marks on public leaderboards and OpenAI accused Moonshot AI of a distillation attack.

The Federal Trade Commission has opened an investigation into OpenAI, Anthropic and other AI companies over the potential dangers their products pose, an agency spokesperson confirmed to CNBC on 30 September. The New York Post first reported the probe.
It is the first official US enforcement action that looks directly at rogue AI agents, according to The Guardian, following a run of incidents first reported in July. The FTC plans to issue formal demands for information and compel testimony from executives, including at the research group Metr, which Anthropic and OpenAI have used for independent incident investigations.
Open weights keep climbing
While the regulators move, the open-weight leaderboards keep reshuffling. BenchLeader's open-weights table, updated in the last day, puts Moonshot AI's Kimi K3 at the top on 66.2, ahead of Zhipu AI's GLM 5.3 on 64.7 and Xiaomi's MiMo-V2.6-Pro on 64.5.
BenchLM.ai ranks the same field differently, placing MiMo-V2.6-Pro first at 75.5 on its BenchAlign v5.8 score, with Qwen3.8 Max and MiMo-V2.6-Flash behind it. The two sites use different scoring contracts and different evidence labels, so the disagreement is about methodology as much as models. LMC Marketcap's directory of 175 open models scores Qwen3.8 27B highest at 71, while The Open Weights, citing Artificial Analysis data refreshed on 1 October, gives GLM-5.3 the top intelligence score at 44.8, ahead of Kimi K3 at 43.6.
Read together, the four rankings agree on one thing: Chinese labs hold most of the top slots.
Nvidia pushed into a different corner of the open catalogue on 29 September. Its Kumo Tabular foundation model for tables, released under the OpenMDW-1.1 licence, comes in three sizes from 28M to 215M parameters, was pretrained only on artificial data, and predicts labels in a single forward pass with no training, tuning or feature engineering, according to the company's Hugging Face post. Nvidia says it ranks first on TabArena, BeyondArena, TALENT and ScoringBench.
Small models, narrow jobs
The week's more interesting open releases are small and specific. Fermion Research published Phonon-2 on 30 September, an English speech recognition model in a 164 MB download that averages 5.21 % word error across the Open ASR Leaderboard's seven English sets, within noise of its 2.5 GB teacher and better on meetings and parliamentary speech. The weights are CC-BY-4.0, derived from Nvidia's Parakeet TDT 0.6B v3.
H Company's Holo4 family, covered by MarkTechPost, puts computer-use weights into two sizes: a 27B dense model and a 35B-A3B mixture of experts. The larger ships under Apache 2.0 for commercial self-hosting; the 27B, which scores 85.2 % on OSWorld at $0.08 per task by H's own figures, is CC BY-NC 4.0, so commercial use runs through the vendor's API. MarkTechPost notes that frontier comparisons used different harnesses and that 480 of AutomationBench's 600 public tasks sit in H's training split.
"Fast and cheap is very easy, you know," TypeSafe CEO Diogo Almeida told TechCrunch. "If you want it really fast and cheap, use dice, right? Intelligence is the hard part."
That quote points at the quiet story underneath the leaderboards: the Jev-style decision-model format, which returns choices and probabilities instead of prose, is spreading fast. OpenAI's Decisions API, announced at DevDay, applies the same idea to its Luna model, TechCrunch reported on 30 September. Bespoke Labs' Nimble, a 9B model built on Qwen3.5-9B, publishes its full training recipe and 2,676 training examples; the repository states plainly that it did not distil from Jev. Ollama added local support for Nimble in version 0.35, and Jevstiller, an open tool covered by The Register on 29 September, claims 98 % agreement with Jev while answering familiar requests on-device in as little as 15 ms.
The accusation, and the shrug
The political temperature rose on 30 September, when OpenAI said it had disrupted an adversarial distillation campaign that ran from 1 July to 28 July, peaking on 24 and 25 July with 16,000 requests from over 4,000 users. OpenAI's blog said the core cluster came from Moonshot AI. The Register noted the irony in a headline and pointed out that Moonshot did not respond to a request for comment.
Moonshot's Kimi K3 sits near the top of several open leaderboards this week. Whether distillation claims change how enterprises pick weights is a separate question, and the token-bill argument is doing more work than any accusation.
Sources
14- 01FTC is investigating OpenAI, Anthropic and other AI companies over product risksEN
- 02US trade regulator opens investigation into AI giants including Anthropic and OpenAIEN
- 03Best open weights AI models ranked (2026)EN
- 04Open-Source LLM Leaderboard 2026: 94 Models RankedEN
- 05Best Open Source AI Models & LLM Leaderboard (2026)EN
- 06LLM Intelligence leaderboardEN
- 07NVIDIA Kumo Tabular: Open Foundation Model for Tabular PredictionEN
- 08Phonon-2: most accurate open speech recognition model in a 164 MB downloadEN
- 09H Company Releases Holo4: Open-Weight Computer-Use ModelsEN
- 10OpenAI's Jev clone could help the frontier lab stop its swarming agentsEN
- 11Nimble: Data, Model, Recipe for an Open Jev (From Bespoke Labs)EN
- 12Ollama now supports Jev-like decision models all locally in 0.35EN
- 13Open source tool distills Jev so you can run it locallyEN
- 14Irony alert: OpenAI whines that Chinese model stole its special IP that it stole from everybody elseEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.