Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Open weights move up the stack: chip design, speech and tabular models ship in one week

OpenAI and Synopsys signed a multi-year deal on 30 September to build GPT-Synopsys, a specialised model for chip design. It was one of at least five open or openly described model releases and partnerships to land in the final days of September.

AI & modelsAnalysisRachel NwosuPublished: 30 September 20268 min readSources 7
Open weights move up the stack: chip design, speech and tabular models ship in one week

The newest development in open model work is not a benchmark score. It is a contract. On 30 September, OpenAI and Synopsys said they had signed a multi-year strategic partnership to build a specialised AI model for chip design called GPT-Synopsys, according to The Decoder. Synopsys sells electronic design automation tools, the software engineers use to lay out chips. OpenAI is licensing those tools for the project.

The stated goal is a model that can "reason about chip design and verification, and to directly operate Synopsys' tools." Both companies say early tests with semiconductor customers are underway. Customer data will not be used for training and will be stored encrypted. The two will market the product together and share revenue. Synopsys CEO Sassine Ghazi says AI could significantly speed up the design process. OpenAI co-founder Greg Brockman frames the deal as a path to better chips and better AI.

That is a closed model inside a commercial pipeline. The same week produced a very different kind of release.

Small models, small files, loud claims

Fermion Research released Phonon-2 on 30 September and calls it the most accurate open speech recognition model under 900 MB. In a 164 MB download, the company says, it matches the accuracy of its 2.5 GB full-precision teacher set for set, beats it on meetings and parliamentary speech, and turns an hour of audio into text in about 20 seconds on a MacBook Air. The weights are on Hugging Face under CC-BY-4.0, the licence of NVIDIA's Parakeet TDT 0.6B v3, from which they derive.

The interesting number is 5.21. Across the Open ASR Leaderboard's seven English sets, Phonon-2 averages 5.21 percent word error. Every open model that scores better is at least 5.8 times its size, according to the company's release. Each encoder weight is stored as one of five learned levels, packed as base-3 digits five to a byte, with one further bit per non-zero weight selecting the magnitude. That comes to roughly 2.1 bits per stored weight. The company says the lead holds under noise, from a quiet room to background noise as loud as the speaker, where it stays ahead of Parakeet Redux.

Independent verification is not in the dossier. Treat the leaderboard position and the noise claim as vendor numbers until someone reproduces them.

Elsewhere the same day, NVIDIA put its tabular foundation model on Hugging Face. NVIDIA Kumo Tabular is an open foundation model for tabular data. It predicts labels of new rows in a single forward pass, with no training, no tuning and no feature engineering, for both classification and regression. It was pretrained only on artificial data and comes in three sizes from 28M to 215M parameters. It runs through NVIDIA's open-source structured-data-models library and is released under the OpenMDW-1.1 licence for commercial use. NVIDIA says it ranks first on four benchmarks: TabArena, BeyondArena, TALENT and ScoringBench.

The architecture details matter less than the framing. Tabular work has been dominated for two decades by gradient-boosted trees. NVIDIA's pitch is that a pretrained transformer can read a labelled table in context and skip the per-task lifecycle of collecting labels, engineering features and tuning hyperparameters. The model uses column, row and in-context attention in the lineage of TabICL and TabPFN, with a length-aware attention temperature so that softmax attention does not dissolve when an inference table is far larger than a training table. Missing values need no imputation.

Open weights from volunteers and from a state institute

Two releases on 30 September came with no company behind them at all, or with a state one. Coop is a small language model pretrained by volunteers, with no server, no funding and no daemon. The training loop runs on donated consumer hardware plus free tiers of Hugging Face and GitHub Actions, according to its GitHub repository. Stage 2 is live: a roughly 145M-parameter decoder-only transformer pretraining from scratch on FineWeb-Edu, with 12 layers, 14 heads, d=896, a 1024-token context and a 32k byte-level BPE vocabulary. Stage 1, a 15M model on TinyStories, finished past its Chinchilla-optimal budget with validation loss falling from 9.01 to 2.8 in six days.

The mechanism is the story. Workers download the current checkpoint, run local AdamW steps on a personal data shard, compute a pseudo-gradient, and submit it as a pull request against a public Hugging Face dataset repository. A stateless GitHub Actions cron job reads the checkpoint and the open inbox pull requests. It is scheduled every five minutes but in practice fires anywhere from minutes to a few hours apart. It drops over-stale submissions, clips and cosine-gates the rest, aggregates them by trimmed mean or geometric median, takes one Nesterov outer step, uploads the new checkpoint, credits contributors in the ledger and closes the processed pull requests.

The repository is candid about evaluation. A single validation loss says nothing, because outer steps move it up as often as down. So the leaderboard fits a slope over the series against tokens rather than outer steps and reports it with a standard error. "Going down" means the slope clears two standard errors. That is a more honest reporting rule than most funded labs manage.

On the same day, the UK AI Security Institute and Meridian Labs published Inspect, an open-source framework for frontier AI evaluations. Inspect covers coding, agentic tasks, reasoning, knowledge, behaviour and multi-modal understanding, and ships with over 200 pre-built evaluations. Its building blocks are datasets, solvers and scorers. It supports tool calling including MCP tools and built-in bash, python, web search and computer tools, and runs arbitrary external agents such as Claude Code, Codex CLI and Gemini CLI. It sandboxes untrusted model code in Docker, Kubernetes, Modal, Proxmox and Vagrant through an extension API. The timing is pointed. Evaluation tooling from a government safety institute lands in the same week that a United States regulator opened an industry-wide investigation into AI agents misbehaving.

Open weights cross into China's chip stack

Deepseek is releasing open-source programming tools for Huawei's AI chips, according to a post on Deepseek's official WeChat channel, with Reuters reporting the open-sourcing. The centrepiece is TileLang, an open-source programming language for AI chips originally developed by researchers at Peking University. Deepseek has used it for about a year and now describes it as its main tool for work on artificial general intelligence, according to The New York Times.

The software includes libraries for computation and for moving data between chips. According to Deepseek, Huawei "fully supported" the work, and the two companies optimised a supernode, a cluster of 128 Ascend 950 chips. Deepseek argues that anyone trying to build an independent software ecosystem for AI chips first needs a universal language that is easy to program but still gets full performance out of the hardware. It says TileLang offers a simpler programming model than CUDA.

That connects to a wider argument about where Nvidia's advantage sits. Research firm SemiAnalysis tested OpenAI's Jalapeno inference chip and called the CUDA moat "potentially dead", because OpenAI gets new models running on its own hardware so quickly. It said Jalapeno beat Nvidia's Blackwell on performance per watt in most scenarios tested. SemiAnalysis attached its own caveats: the tests covered relatively easy-to-optimise scenarios of about 8,000 input tokens and 1,000 output tokens, and it has not yet run AgentX, the benchmark for multistep agent tasks. In August, SemiAnalysis found Nvidia well ahead on that benchmark. It said Nvidia would still be cheaper per token than AMD even if AMD gave its hardware away, because the advantage sits in the software that links many chips into one system. Huawei's chips were not part of the AgentX comparison.

There is a domestic pressure point too. Huawei admits it cannot keep up with demand at home, so it plans to sell fewer chips abroad. Referring to US export controls, Huawei's current rotating chairman Eric Xu said the company cannot accept a future that hinges on whether others are willing to sell chips to China.

What to watch

The pattern across these releases is that open weights are no longer confined to chatbot checkpoints. They are arriving as speech encoders small enough for a laptop, tabular predictors aimed at enterprise tables, evaluation harnesses from a state safety body, volunteer training loops that run on free tiers, and now programming tools for a non-Nvidia accelerator stack.

Two things remain unresolved. First, the benchmark claims for Phonon-2 and Kumo Tabular are vendor or maintainer claims in this dossier, and none of the three model releases has an independent reproduction attached. Second, the regulatory picture is moving in the opposite direction from the releases. The FTC's investigation into Anthropic, OpenAI and Metr is the first official US enforcement action touching rogue AI agents, as The Guardian reported on 30 September.

Open weights do not settle any of that. They just move the argument onto hardware more people can run.

Comments 0

Sources

7
  1. 01OpenAI and Synopsys team up to build an AI model that designs chips like a seasoned engineerEN
  2. 02Phonon-2: most accurate open speech recognition model in a 164 MB downloadEN
  3. 03NVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular PredictionEN
  4. 04Coop: A small language model pretrained by volunteersEN
  5. 05Inspect: An open-source framework for large language model evaluationsEN
  6. 06China's AI industry closes ranks as Deepseek ships open-source software for Huawei's Ascend chipsEN
  7. 07US trade regulator opens investigation into AI giants including Anthropic and OpenAIEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.