Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Open weights close the gap: new leaderboards, new hardware, and a chip deal

On 1 October, the Artificial Analysis Intelligence Index as republished by The Open Weights put Zhipu AI's GLM-5.3 at the top of its 20-model open-weight table with a score of 44.8, ahead of Moonshot AI's Kimi K3 at 43.6, as fresh releases keep arriving.

AI & modelsAnalysisGrace OkonkwoPublished: 30 September 20264 min readSources 14
Open weights close the gap: new leaderboards, new hardware, and a chip deal

On 1 October, the Artificial Analysis Intelligence Index as republished by The Open Weights put Zhipu AI's GLM-5.3 at the top of its 20-model open-weight table with a score of 44.8, ahead of Moonshot AI's Kimi K3 at 43.6, as fresh releases keep arriving.

The rankings disagree with each other, which is worth saying out loud. BenchLeader's open-weights board, updated 12 hours before publication, puts Moonshot AI's Kimi K3 first at 66.2, with Zhipu's GLM 5.3 max second at 64.7 and Xiaomi's MiMo-V2.6-Pro third at 64.5. BenchLM.ai's 94-model board instead lists MiMo-V2.6-Pro first with a BenchAlign v5.8 score of 75.5, with Qwen3.8 Max and MiMo-V2.6-Flash behind it. LMMarketCap, which published an update at 06:00 on 1 October, ranks Alibaba's Qwen3.8 27B first among 175 open models, with Google's Gemma 4 26B A4B and Gemma 4 31B joint second.

Different panels, different winners. The one thing the four boards share is that the top rows are Chinese labs and Google, with Nvidia, Meta and Mistral further down.

The last 72 hours also produced several new open models that are not on those boards yet. Nvidia published Kumo Tabular on 29 September, a foundation model for tabular classification and regression that comes in three sizes from 28M to 215M parameters, was pretrained only on artificial data and is released under the OpenMDW-1.1 license for commercial use; the company says it ranks first on TabArena, BeyondArena, TALENT and ScoringBench, and it needs no training or feature engineering per task.

Fermion Research released Phonon-2 on 30 September, an English speech recognition model in a 164 MB download whose encoder stores each weight in about 2.1 bits. The company says it averages 5.21 percent word error across the Open ASR Leaderboard's seven English sets, that every open model scoring better is at least 5.8 times larger, and that an hour of audio transcribes in about 20 seconds on a MacBook Air. The weights are CC-BY-4.0, derived from Nvidia's Parakeet TDT 0.6B v3. Bespoke Labs published Nimble on 30 September, a 9B model built on Qwen3.5-9B that returns typed decisions with probabilities instead of free text; its README says the training and held-out examples were published on 20 September after being left out of the first release.

Two open projects are testing whether the training loop itself can be distributed. The Commonsense AI coop repository, updated 30 September, describes a roughly 145M-parameter model pretraining from scratch on FineWeb-Edu where volunteers run local AdamW steps on donated hardware, submit pseudo-gradients as Hugging Face pull requests, and a stateless GitHub Actions cron every five minutes aggregates them with a trimmed mean or geometric median. Stage 1, a 15M model on TinyStories, completed past its Chinchilla-optimal budget. Separately, PSSA, a non-transformer language model written in Rust with no ML framework underneath, claims it learns faster than a transformer at matched parameters and generates about twelve times quicker on the same CPU.

The commercial side moved too. OpenAI and Synopsys said on 30 September that they had signed a multi-year partnership to build GPT-Synopsys, a model for chip design and verification that would operate Synopsys' EDA tools directly, running on OpenAI infrastructure, with customer data not used for training and revenue shared between the two. The Register reported on 30 September that OpenAI accused individuals associated with Moonshot AI of a distillation campaign starting 1 July, with spikes on 24 and 25 July of 16,000 requests from over 4,000 users and related activity across more than 15,000 accounts, disrupted on 28 July. Moonshot did not respond to The Register's request for comment.

Regulators are circling the same labs. The FTC has opened an industry-wide investigation into OpenAI, Anthropic and others, CNBC confirmed on 30 September, and The Guardian reported the same day that the agency plans formal demands for information and testimony, including from the research group Metr. Neither OpenAI nor Anthropic commented to CNBC; the New York Post reported the probe first. On 29 September a nonprofit, Legal Advocates for Safe Science & Technology, sued OpenAI in San Francisco County Superior Court over the July Hugging Face hack, seeking an injunction rather than damages, according to Ars Technica, which quoted OpenAI calling the suit "completely without merit".

For anyone choosing an open model today, the practical lesson from this week is that no single leaderboard settles the question: the boards rank different sets of models, and a 164 MB speech model or a 215M tabular model can matter more to a deployment than whoever sits at number one.

Comments 0

Sources

14
  1. 01LLM Intelligence leaderboard - The Open WeightsEN
  2. 02Best open weights AI models ranked (2026) - BenchLeaderEN
  3. 03Open-Source LLM Leaderboard 2026: 94 Models Ranked - BenchLM.aiEN
  4. 04Best Open Source AI Models & LLM Leaderboard (2026) - LMMarketCapEN
  5. 05Nvidia Kumo Tabular: Open Foundation Model for Tabular PredictionEN
  6. 06Phonon-2: most accurate open speech recognition model in a 164 MB downloadEN
  7. 07Nimble: Data, Model, Recipe for an Open Jev (From Bespoke Labs)EN
  8. 08Coop: A small language model pretrained by volunteersEN
  9. 09PSSA: A non-transformer language model written from scratch in RustEN
  10. 10OpenAI and Synopsys team up to build an AI model that designs chips like a seasoned engineerEN
  11. 11Irony alert: OpenAI whines that Chinese model stole its special IP that it stole from everybody elseEN
  12. 12FTC is investigating OpenAI, Anthropic and other AI companies over product risksEN
  13. 13US trade regulator opens investigation into AI giantsEN
  14. 14"An AI did it" is no defense, says nonprofit suing OpenAI over Hugging Face hackEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.