Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Model trackers multiply as release dates and training cutoffs diverge

A tracker published on 16 September lists 20 current AI models across 8 labs. It shows the gap between the day a model ships and the day it stopped reading. Only 10 of the 20 carry a training cutoff their lab actually publishes.

AI & modelsAnalysisRachel NwosuPublished: 27 September 20263 min readSources 3
Model trackers multiply as release dates and training cutoffs diverge

The page, stale.jock.pl, keeps two dates per model: the release date, when the lab shipped it, and the training cutoff, when it stopped reading. Counters tick upward live from both. The author puts it bluntly: a model can ship in September and still stop reading in April, which leaves it five months behind on launch day.

The list is concrete. Llama 4 shipped on Apr 5, 2025 with an August 2024 cutoff. Claude Haiku 4.5 shipped on Oct 15, 2025 and stops at July 2025. Gemini 3.1 Pro shipped on Feb 19, 2026 with a January 2025 cutoff, thirteen months behind. GPT-6 Astra shipped on Sep 3, 2026 and stops at Apr 30, 2026. GPT-6 Sol, from Sep 22, 2026, stops at Apr 20, 2026; GPT-6 Luna, same day, stops at May 18, 2026. Claude Opus 5.5, also Sep 22, stops at June 2026.

Half the shelf has no published cutoff

Ten models carry a blank. Five of the eight labs, Anthropic, Google DeepMind, Meta, OpenAI and xAI, have a published cutoff for at least one model on the page. The blanks include the entire Mistral line shown (Large 3, Small 4, Medium 3.5), Qwen3.8-Max and Qwen3.8-Flash from Alibaba, Meta's Muse Glimmer and Muse Spark 1.3, DeepSeek V4-Pro and V4.1-Flash, and Gemini 3.8 Flash. The page is careful about what a blank means: the vendor sources it checked did not establish a cutoff, which is not proof the lab never published one.

The author also ran a search test: over 2,000 calls across 16 models, each with a web search tool. Frontier models decided correctly almost every time, the page says. Weaker ones stated settled facts that had changed without checking, and searched the web for things like the boiling point of water. In the same test, five models named a dead man as the king of Norway.

The search also has to be triggered by the model, using the same weights that hold the stale fact, so the misses land exactly where the model feels most certain.

That is the argument for the export. The page offers models.json with release date, published cutoff and a source link per model. It suggests pasting a few lines into the AGENTS.md or CLAUDE.md an agent already reads. The instruction is to fetch the file before naming any model, version or date as current, and to treat anything with a date attached as unverified until checked. The file is regenerated when the page is.

Benchmarks move, the dates do not

This is not the only attempt to police model facts. A separate Show HN project, Cactus Needle 3, published on 18 September, pitches 8 to 29 MB automation models for tiny devices and claims a 4-layer subnetwork can match DeepSeek V4 Flash when tuned for one epoch. CUA-S1, posted on 19 September, is an early source-only research release of small decision models for computer use, with weights hosted separately on Hugging Face and MIT-licensed source.

None of these projects settles the benchmark question. A tracker tells you when a model stopped reading. It does not tell you whether the model is any good. The distinction matters because release announcements and benchmark claims arrive on different clocks. The cutoff date is the one piece of metadata a lab can publish without arguing about methodology.

The page's own advice is circular, and it admits it: ask the model for its training cutoff directly. A well behaved model answers or says it is not sure. One that invents a confident date has told you something useful about itself. Then check the answer, because a model is a poor source on models.

Comments 0

Sources

3
  1. 01How stale is your AI? Release age and training cutoff for 20 modelsEN
  2. 02Needle 3 - 8-29 MB foundation model for tiny devicesEN
  3. 03trycua/cua: Scale computer-use 2.0EN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.