Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

A release date is not a training cutoff: tracking how stale 20 AI models already are

Ten of 20 current AI models carry a training cutoff their lab actually publishes, according to a tracker published on 16 September. The gap between shipping date and cutoff is the number most launch coverage never mentions.

AI & modelsAnalysisGrace OkonkwoPublished: 27 September 20263 min readSources 1
A release date is not a training cutoff: tracking how stale 20 AI models already are

The page at stale.jock.pl holds two dates for each model: the day the lab shipped it, and the day it stopped reading. Both counters tick upward live. The second date decides what the model actually knows.

Take GPT-6 Astra, released by OpenAI on 3 September 2026. Its training data stops at 30 April 2026. On launch day the model was already four months behind on anything that happened after that date. It will keep falling further behind, because nothing in the weights changes when the world does. The tracker lists the same pattern across the shelf. Llama 4 from Meta shipped 5 April 2025 with an August 2024 cutoff. Claude Sonnet 5 from Anthropic shipped 30 June 2026 with a January 2026 cutoff. GPT-6 Sol shipped 22 September 2026 with a 20 April 2026 cutoff. Each name on the page links to the lab document the date came from, according to the site.

Half the list is blank.

Ten of the 20 models carry a published cutoff. The other ten, including Mistral Large 3, Qwen3.8-Max, DeepSeek V4-Pro and Muse Glimmer, show no cutoff because the checked vendor sources did not establish one, the page says. That is not proof a lab never published one. It is proof the tracker could not find it, which is a different and more awkward fact for anyone trying to compare models on freshness.

Search does not fix the gap

The obvious objection is that models search the web now, so the cutoff matters less. The tracker's author rejects that. Search reads a few pages for one answer and forgets them, the page argues, so a new chat starts from April again. Worse, the model itself has to trigger the search, using the same weights that hold the stale fact. The misses land exactly where the model sounds most certain. The page cites a test of over 2,000 calls across 16 models, each given a web search tool. Frontier models decided correctly almost every time. Weaker ones stated facts that had changed without checking, and searched the web for things like the boiling point of water. In one example, five models named a dead man as the king of Norway.

There is an uncomfortable methodological point buried here. A model asked for its own training cutoff is a poor witness, the page says, because one that invents a confident date has told you something useful about itself. The recommended check is external.

The tracker ships the same data as models.json, with release date, published cutoff and a source link per model. It suggests pasting a few lines into an agent's AGENTS.md or CLAUDE.md so the agent fetches the file before naming any model or version as current. The file is regenerated when the page is, so the agent reads today's dates rather than the ones baked into its weights. The 20-model list spans 8 labs, and 5 of those labs have a published cutoff for at least one model on the page: Anthropic, Google DeepMind, Meta, OpenAI and xAI.

For benchmark watchers, the practical consequence is simple. A leaderboard comparing models released weeks apart is also comparing models trained months apart, and the delta is not in the score.

Comments 0

Sources

1
  1. 01How stale is your AI? Release age and training cutoff for 20 modelsEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.