Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

The release date is not the story: training cutoffs and adoption data redefine AI model benchmarks

Ten of the 20 current AI models listed on stale.jock.pl carry a training cutoff their own lab publishes, and on OpenRouter Chinese models went from 6%-13% of tokens in February 2026 to 57%-67% in the week of Sept. 14. Neither number appears in a standard benchmark table.

AI & modelsAnalysisGrace OkonkwoPublished: 27 September 20265 min readSources 3
The release date is not the story: training cutoffs and adoption data redefine AI model benchmarks

A benchmark table tells you what a model scored. It does not tell you when the model stopped reading. That gap, between the release date on a launch post and the training cutoff buried in a model card, is now one of the more useful numbers a buyer can look up. It is also the premise of a small tool published on 16 September at stale.jock.pl under the title "How stale is your AI?"

The page lists 20 current models across 8 labs and pairs each release date with a training cutoff, counting upward from both live. Ten of the 20 have a cutoff the lab actually publishes. Five of the eight labs, Anthropic, Google DeepMind, Meta, OpenAI and xAI, have a published cutoff for at least one model on the list. The rest are marked "Not established". The page is careful to say this is not proof the lab never published one, only that the checked vendor sources did not turn one up.

Some of the gaps are large.

GPT-6 Astra, released by OpenAI on 3 September 2026, has a training cutoff of 30 April 2026. It was already four months behind the day it shipped. Gemini 3.1 Pro, released 19 February 2026, stops at January 2025, a gap of more than a year. Claude Sonnet 5, out 30 June 2026, has a January 2026 cutoff; Claude Opus 5.5, out 22 September 2026, has June 2026. Llama 4 from Meta, released 5 April 2025, carries an August 2024 cutoff.

The page's central argument is that a search tool does not fix this. When a model searches the web it reads a few pages, uses them for that one answer and forgets, the site says. The model itself has to trigger the search, using the same weights that hold the stale fact, so misses cluster where the model is most confident. The author reports running 16 models through a web search tool over 2,000 calls. Frontier models decided to search correctly almost every time. Weaker ones answered settled questions from memory after the answer had changed, and searched the web for things like the boiling point of water. In one example, five of the 16 named a dead man as king of Norway.

None of this is peer reviewed. The page is a Show HN project, the numbers come from the author's own test harness, and the model list is a snapshot rather than a census. But the underlying point does not depend on the experiment holding up. A model that launched in September with an April cutoff is wrong about September in a way that a benchmark score will not show you.

There is a second, harder-to-ignore signal in the same week, and it is about which models people actually run. CNBC reported on 26 September that Chinese models went from a small share of usage to a majority on two developer gateways. On OpenRouter they accounted for 57% to 67% of tokens in the week of 14 September, up from 6% to 13% in February. On Vercel their share rose to 55% in August from 11% in January.

OpenRouter's figures cover companies in the U.S., Europe and what it defines as the "Global South", 82 countries across Central and South America, Africa and Asia. Vercel did not give a geographical breakdown. The same report notes that U.S. frontier models still attract more overall spending, so token share and revenue share point in different directions.

Peter Walker, head of insights at OpenRouter, told CNBC that Chinese open source models released this year "can credibly perform in advanced agentic use cases, especially in regards to coding, in a way that was just not true in late 2025". They are also "incredibly cost-effective compared to most models from American labs", he said. Harpreet Arora, head of agentic infrastructure at Vercel, put the mechanism plainly: once a model meets the quality bar for a job, the price difference becomes compelling.

Washington is paying attention. Two U.S. House committees are investigating the rising adoption of Chinese models. CNBC reports concern about remote access to Nvidia chips through overseas data centers and about distillation, where newer models imitate older ones. Daniel Remler of the Center for a New American Security told CNBC the "ultimate concern is that the integration of Chinese AI models pulls countries into a Chinese technology sphere of influence that hardens into geopolitical alignment".

Put the two stories side by side and the benchmark question changes shape.

A leaderboard measures a model against a fixed set of tasks on the day it is tested. It does not measure how far the model's knowledge has drifted from the present, and it does not measure whether anyone can afford to run it at volume. The stale.jock.pl page argues that models are poor sources on themselves. It suggests pointing an agent at a machine-readable models.json file before it names any model, version or date as current. That is a small fix. The larger problem is that the industry's default evidence, the benchmark table, is silent on both freshness and price.

Nvidia is making a related bet, according to Rest of World. The company spent $7 billion last month for a stake in U.S. coding startup Poolside and a licence to its model-building tools, and has agreed to pay almost $13 billion for Hugging Face. Its own free model, expected to be called Nemotron 4, is due by the end of the year. The report says earlier U.S. open models have lagged China's free models, which were downloaded roughly 2 billion times this year according to an August Hugging Face report. Nvidia's spending is meant to make Nemotron the base governments pick instead of Meta's or Alibaba's.

The UAE is the test case. Its flagship data center will run on 400,000 Nvidia chips, and G42, the state-backed AI company, joined Nvidia's open-model alliance in July. The Technology Innovation Institute, which built Falcon, is taking the other side. "As a model-building lab, we believe that maintaining diversity in AI perspectives, and preserving full technological sovereignty, requires a strong, homegrown foundation," TII chief researcher Hakim Hacid told Rest of World.

Neither position is settled by a benchmark. A country or a company choosing a base model is choosing a supply chain, a cost curve and a refresh cadence. The published cutoff is one of the few hard numbers available about that last part.

Comments 0

Sources

3
  1. 01Show HN: How stale is your AI? Release age and training cutoff for 20 modelsEN
  2. 02Chinese AI models surge in global popularity — and Washington is worriedEN
  3. 03Nvidia's free AI model could push the UAE closer to the U.S.EN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.