Soofi S benchmark row dropped after GPQA test questions found in its training data
A German research consortium has removed the science benchmark GPQA from its evaluation of Soofi S, an open 30B model, after discovering that rephrased test questions had leaked into the disclosed training data, according to its pretraining report.

The consortium, coordinated by the KI Bundesverband, published version 3.0 of the pretraining report and documented the contamination itself. The affected dataset was QA-base. It was meant to hold only rephrased training splits from 25 standard benchmarks: practice questions that teach a model test formats, not the real test items. It also held rephrased questions from the GPQA test set, including the harder GPQA Diamond variant, in English and in machine-translated German.
The cause was a labelling quirk, not a deliberate shortcut.
On Hugging Face, GPQA ships without a separate training split. All of its test material sits under the default label "train." The consortium's data pipeline selected content by split name, so test questions were swept in with the practice items. An audit triggered by the discovery turned up three more benchmarks with the same pattern: TruthfulQA, BLiMP and Inverse Scaling. None of those three are used in the Soofi S evaluation.
A gain that looked normal
Nobody caught the problem during training. The model improved steadily on the contaminated GPQA tasks, climbing from 32.3 to 43.4 points. The authors say that gain fell inside the normal range and looked no different from progress on other tests. There was no spike that could have served as a warning.
The fix is procedural. The team now cross-checks its training data against all test questions automatically, rather than relying on split labels. The report describes those labels as often misleading.
In response to the leak, the consortium dropped GPQA entirely from its evaluation. It also recalculated overall results for all 16 compared models, so the field would be scored on the same basis. According to the authors, the ranking order did not change.
Project participant Nicolas Flores Herr of Fraunhofer IAIS points to the incident as evidence that the open approach works. The community could only find the problem, he argues, because the data was public. The team says it has made roughly 152,000 individual results available for verification. A separate evaluation by consortium partner Ellamind confirmed the results. That run used deliberately withheld test data the model was guaranteed not to have seen in training. The consortium says the affected portion is a tiny fraction of the total corpus.
The timing is awkward for benchmark-based claims generally. Separate coverage in recent days, including work from Stanford HAI, has questioned how well current AI tests grade the systems they are aimed at.
What Soofi S actually is
Soofi S 30B-A3B is a mixture-of-experts model. It carries 31.6 billion parameters in total but activates about 3.2 billion per generated token. That puts its compute cost closer to a 3B model than to a conventional 30B one. The consortium adopted the architecture of Nvidia's Nemotron 3 Nano without modification: a hybrid design that combines Mamba-2 layers with standard attention layers.
The practical difference is memory behaviour. In a conventional transformer, the KV cache holding previous tokens grows linearly with context length, and reloading it becomes a bottleneck under long inputs and many parallel requests. Only 6 of Soofi S's 52 layers maintain such a cache at all. At a context length of 40,000 tokens with 32 parallel requests, the consortium reports Soofi S generating roughly eight times more tokens per second per GPU than dense models in the 14 to 24 billion parameter range. Throughput stays nearly flat from 4,000 to 256,000 tokens, while conventional models drop off. According to the report, the only model showing similar behaviour in the measurements is Alibaba's Qwen3.5 35B-A3B, which also uses a hybrid architecture.
Training ran entirely on Deutsche Telekom's Industrial AI Cloud in Munich, making Soofi S one of the first large language models built on that infrastructure. The training mix is deliberately weighted toward German.
On the consortium's own benchmarks, the model beats other fully open systems, including OLMo 3 32B and Apertus 70B, across German, English and programming tasks. The fine-tuned instruct and reasoning variants are in beta testing and are expected to ship under a permissive license in the coming weeks. The larger Soofi L model is already in training.
The overtrained argument
After launch, critics argued Soofi S was heavily overtrained by the standards of the Chinchilla scaling laws. Google DeepMind published those laws in 2022 to describe how to balance model size against training data for a fixed compute budget. The sweet spot they identified was roughly 20 tokens per parameter. Soofi S blows past that. With about 27 trillion tokens and 30 billion parameters, it lands at several hundred to one. Count only the 3.2 billion parameters active per token and the ratio reaches several thousand to one.
Michael Fromm, part of the project's technical leadership, rejects that reading. He argues the old rules do not carry over to mixture-of-experts architectures. "There's new research showing that the old scaling laws from dense models no longer apply to MoE architectures," Fromm said. The reason, he says, comes down to construction: individual experts benefit from seeing the same documents, so repeated data in a large, high-quality set is less of a problem than it would be for a dense model. As a comparison point, he cites Nvidia, which trained its own models on up to 25 trillion tokens.
The contaminated data was test material, not practice material, and the pipeline could not tell the difference because of a label.
That is the uncomfortable part of the episode. The leak was not a case of a lab quietly training on a test set. It was a naming convention on a public dataset, applied at scale by an automated pipeline, and it slipped past everyone until the data was out in the open and someone went looking. The consortium's answer, cross-checking every training item against every test question, is the kind of control that costs compute and time and that few teams describe publicly.
The broader context for this release is a benchmark culture under strain. A tracker published at stale.jock.pl in September lists release dates and training cutoffs for 20 current models across 8 labs. It notes that only 10 of the 20 carry a cutoff the lab actually publishes. Its authors argue a model can ship in September and still have stopped reading in April, and that search tools paper over the gap rather than closing it. Measured over 2,000 calls across 16 models, each given a web search tool, they report that frontier models decided correctly almost every time, while weaker ones answered settled questions from memory after the answer had changed.
Benchmarks and model cards are the industry's shared vocabulary for what a system can do. Soofi S shows both halves of that: a documented leak handled in public, and a set of throughput and benchmark numbers that now depend on a corrected evaluation. The instruct and reasoning variants, and the recalculated comparison table, will be the next things to check.
Sources
2- 01German AI consortium releases Soofi S, an open 30B model that tops benchmarks in both English and GermanEN
- 02How stale is your AI? Release age and training cutoff for 20 modelsEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.