AI Model Releases Are Accelerating, but Benchmarks Cannot Keep Up
Ten of the 20 current models tracked by stale.jock.pl ship without a published training cutoff, while Chinese models took 57% to 67% of tokens on OpenRouter in the week of Sept. 14.

Six model launches landed between Sept. 1 and Sept. 22, 2026, according to the model tracker at stale.jock.pl. The tracker lists release dates and training cutoffs for 20 current models across eight labs. OpenAI alone shipped three on Sept. 22: GPT-6 Sol, GPT-6 Luna and Claude Opus 5.5 from Anthropic on the same day, per the tracker's table and recent headlines. The release calendar for frontier AI models is now measured in days, not quarters.
But the instrument used to judge those models is not keeping pace. Training cutoffs, the dates on which a model stopped reading, are published for only half the models on that list. Five of the eight labs, Anthropic, Google DeepMind, Meta, OpenAI and xAI, publish a cutoff for at least one model, the tracker says. A blank entry means the checked vendor sources did not establish a cutoff, not proof that none exists.
The gap between shipping and knowing
The tracker's own framing is blunt: two dates decide how current a model really is. The release date is when the lab shipped it; the training cutoff is when it stopped reading. GPT-6 Astra, released Sept. 3, 2026, has a cutoff of Apr. 30, 2026, according to the page, meaning it was roughly four months behind on launch day. Claude Fable 5.1, released Sept. 1, 2026, carries a June 2026 cutoff. Gemini 3.1 Pro, released Feb. 19, 2026, stopped reading in January 2025, more than a year before it shipped.
A model can ship in September and still stop reading in April, which means it is five months behind on the day it launches.
That gap matters more as release cycles compress. Mistral AI shipped Mistral Large 3 on Dec. 2, 2025, Mistral Small 4 on Mar. 16, 2026 and Mistral Medium 3.5 on Apr. 28, 2026, none with a published cutoff, according to the tracker. Alibaba's Qwen3.8-Max (Aug. 3, 2026), Qwen3.8-Flash (Aug. 26, 2026) and DeepSeek's V4-Pro (Aug. 13, 2026) and V4.1-Flash (Sept. 10, 2026) are also listed without published cutoffs. The same is true of Meta's Muse Glimmer and Muse Spark 1.3, and Google DeepMind's Gemini 3.8 Flash.
Search tools do not close the gap, the page argues. A model that searches the web reads a few pages, uses them in one answer and forgets; open a new chat and it is back to its cutoff. Worse, the search has to be triggered by the model itself, using the same weights that hold the stale fact, so errors land where the model feels most certain. The page cites a measurement of over 2,000 calls across 16 models, each given a web search tool: frontier models decided correctly almost every time, while weaker ones stated settled facts that had changed and searched the web for things like the boiling point of water.
Usage is shifting faster than benchmarks
While labs race to ship, the usage data tells a different story about which models developers actually reach for. Chinese AI models went from a small share of usage to a majority on two major developer platforms in 2026, according to CNBC, which cited usage data shared with the network. On OpenRouter, Chinese models accounted for 57% to 67% of tokens used in the week of Sept. 14, up from 6% to 13% in February. On Vercel, their share rose to 55% in August from 11% in January.
The adoption is concentrated in what OpenRouter defines as the Global South, 82 countries across Central and South America, Africa and Asia. More than two-thirds of the tokens those companies use go to Chinese models, CNBC reported. About half the tokens on OpenRouter are used by companies in the U.S.
Price is the driver, according to people CNBC spoke to. Peter Walker, head of insights at OpenRouter, told the network that Chinese open source models released this year "can credibly perform in advanced agentic use cases, especially in regards to coding, in a way that was just not true in late 2025." They are also "incredibly cost-effective compared to most models from American labs," he said. Harpreet Arora, head of agentic infrastructure at Vercel, told CNBC that once a model meets the quality bar for a job, the price difference becomes compelling, though companies still want frontier U.S. models for more complicated tasks.
Earlier this week, OpenAI and Anthropic both announced new, cheaper models, CNBC reported. Dianne Penn, head of product management, research and labs at Anthropic, told the network the company was trying to make its models' answers "more efficient, so it uses less tokens depending on your effort setting."
Washington is watching the token counts
The shift is drawing scrutiny in Washington, where two U.S. House Committees are investigating the impact of rising adoption of Chinese models, according to CNBC. The U.S. has restricted Chinese AI companies from buying the most advanced chips through export controls; Washington is concerned about them accessing Nvidia chips remotely through overseas data centers and about "distillation," where new models mimic older, more established ones.
Chinese AI represents "real economic and security risks for the United States," Daniel Remler, a senior fellow in the technology and national security program at CNAS, told CNBC. "The ultimate concern is that the integration of Chinese AI models pulls countries into a Chinese technology sphere of influence that hardens into geopolitical alignment," he said. Remler added that Southeast Asia in particular may see significant uptake, and that "anywhere from Lagos to São Paulo to Jakarta where entrepreneurs and governments are looking for cheap, open models, will look first to Chinese AI."
Nvidia is trying to change that calculus with software rather than chips. The company is preparing a free model, expected to be called Nemotron 4, due by the end of the year, aimed at buyers already running Nvidia hardware, Marc Einstein, an analyst at Counterpoint Research, told Rest of World. Few countries fit that description better than the UAE, whose flagship data center will run on 400,000 Nvidia chips, the publication reported. Nvidia has also agreed to pay almost $13 billion to buy Hugging Face, the platform where developers share and download models, according to Rest of World.
The context for that spending: earlier versions of U.S. AI models have lagged well behind China's free models, which were downloaded roughly 2 billion times this year, far ahead of any Western rival, according to an August report by Hugging Face cited by Rest of World. Governments building their own models mostly start from a free one made by Meta or Alibaba, and Einstein said Nvidia's spending is meant to make Nemotron the base they pick instead.
What the benchmarks miss
None of this resolves the underlying measurement problem. A model's benchmark scores say nothing about whether it knows what happened after its cutoff. The stale.jock.pl page suggests a low-tech check: ask a model directly for its training cutoff date. A well-behaved model answers or says it is not sure; one that invents a confident date has told you something useful about itself, the page argues, adding that the answer should be checked against the tracker because a model is a poor source on models.
The page also ships a machine-readable export, models.json, carrying release date, published cutoff and a source link for all 20 models, and suggests pointing an agent at it through the AGENTS.md or CLAUDE.md instructions file the agent already reads. The export is regenerated when the page is, so an agent reads today's dates rather than the ones baked into its weights.
That is a stopgap, not a fix. Labs control what they disclose, and half the current models on the tracker still carry no published cutoff. Until that changes, anyone comparing models on benchmarks alone is grading a snapshot that was already out of date when it shipped.
Sources
3- 01Show HN: How Stale Is Your AI? Release age and training cutoff for 20 modelsEN
- 02Chinese AI models surge in global popularity — and Washington is worriedEN
- 03Nvidia's free AI model could push the UAE closer to the U.S.EN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.