Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Model rankings that the evidence cannot support

Four independent papers from recent weeks show that tables of language model scores look more decisive than the data allows. The problem also touches research done with the help of AI.

ScienceAnalysisSofia MarchettiPublished: 25 September 20268 min readSources 5
Model rankings that the evidence cannot support

A leaderboard table is the usual way machine learning reports progress: rows of models, columns of tests, a higher position meaning better. Recent papers ask how much such a table really tells you once you probe the stability of the measurement. The answer is uncomfortable.

The same answer, a different structure

One audit looked at how a model infers prompt structure. Its author repeated the measurement on eight open variants from five families, from 8 to 675 billion parameters, with caching disabled and 293 preserved intermediate representations. The phenomenon itself turned out to be unstable. Identical calls did not reproduce an identical structure. The average agreement of the node set ranged from 0.39 to 0.96, and in 72 percent of cases a prompt-model pair was never consistent across the whole set. After applying a clustered bootstrap, only the bottom of the ranking held up: the two least reproducible models kept their position in 99 and 86 percent of repetitions, the four middle ones in 27 to 48 percent, and the two best in 68 percent. So the table reliably identifies the worst model, not the best. Two equally sensible rules for pooling repeated campaigns change four of the eight rows and move the study's headline figure by 7 percentage points. Four of the eight endpoints were withdrawn within ten weeks of the measurement.

Selection and statistics

The second paper asks how many hidden model variants may sit behind a published advantage. For a fixed model family, the authors derive a sensitivity curve describing the maximum number of such variants as a function of a lower bound on the correlation within the family. An audit of 394 claims about adjacent positions on the Open LLM Leaderboard showed that 391 of them have no statistical support even before selection is taken into account.

The third paper takes on reliability metrics for model judges borrowed from classical test theory. At a fixed measured judge error rate of 4.72 percent, the KR-20 coefficient ranges from 0.01 to 0.68 depending on how the task bank is redesigned. The dependency index is sometimes confused with the probability of classification, and it differs from it by 0.25 to 0.43 points. The authors' conclusion is a generous one: this is not a review of specific publications but a warning, because the numbers travel further, into deployment decisions and disclosure documents.

Data, not algorithms

The fourth audit covers seven public datasets for predicting educational outcomes. Three passed all four checks before modelling. The other four either failed tests of generalisation between groups or lacked the metadata needed to run them. The sharpest example: on the UCI Student dataset, R² fell from 0.242 under a random split to -0.097 under a group split, and on Higher Ed from 0.041 to -8.79. Adding model complexity did not remove this. On fragile data, model ensembles amplified the instability.

There is also an ethical and editorial thread. In a piece from Guangming Daily published by ScienceNet, five researchers describe how AI is entering the whole research process: from generating ideas, through data processing, to writing the paper. Questions arise about who owns the idea, about smoothing results, and about responsibility that cannot be shifted onto a tool. The common denominator here is the same as in the rankings: a number or a table does not replace a verification procedure.

Comments 0

Sources

5
  1. 01How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt StructureEN
  2. 02How Sensitive Are LLM Leaderboard Claims to Hidden Model Selection?EN
  3. 03Three Ways Classical Test Theory Misleads for LLM JudgesEN
  4. 04Limited Structural Reliability in Public Educational Prediction BenchmarksEN
  5. 05科学网/光明日报: AI时代,我们需要什么样的学术规范ZH

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Sofia Marchetti

Sofia Marchetti

Science and health

Sofia Marchetti covers science and health for FLASH24, working from primary literature, preprints, and agency data rather than press releases. She checks sample sizes, confidence intervals, and whether a study's numbers match its abstract before filing. She interviews researchers and clinicians directly, tracks conference calendars for embargoed results, and compares new findings with earlier trials on the same question. Outside the newsroom she works on materials physics and stargazes through a home telescope, which keeps her close to how measurement error actually behaves. She does not publish a health claim without a named source and the underlying data.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.