Small models on device: 91% of token entropy signals unusable, paper says
Token-level confidence signals are effectively blind in small language models with fewer than 3 billion parameters, according to an arXiv preprint submitted on 21 July. Mean token entropy sat near zero in 91% of dataset-model combinations, whether the answer was right or not.

Prashant Mudgal tested seven approaches across 7 model pairs and 5 standard NLU benchmarks, all on consumer hardware. His conclusion is blunt: token-level entropy early stopping does not work at this scale. Semantic entropy does. It generates multiple samples and clusters answers by meaning, and it recovers a usable signal. Routing uncertain queries to a larger expert model lifted accuracy by up to 50 percentage points.
Cross-family routing averaged +22.0% against +6.8% for same-family pairs. The example given is SmolLM 360M handing off to Phi-3.5-mini. The paper's framing is that the value of entropy methods in small models is not compute saved but compute allocated.
Why on-device work keeps growing
On 27 September, Nvidia released Nemotron 3 Diarization, a roughly 100 million parameter model that identifies who is speaking in a conversation, with weights free to download, as reported by The Decoder. It separates up to eight speakers and flags overlapping speech. On the Diarization-Bench from VoiceArena it currently sits first with a 14.72 percent error rate, ahead of the next best system at 19.3 percent. Paired with a speech recognition system such as Parakeet, it produces transcripts with anonymous speaker labels. Error rates rise with more participants, background noise or reverb.
Latency is a design trade-off, not a free win. The audio buffer can be set to four levels from 30.4 seconds down to 0.32 seconds, and shorter buffers generally reduce accuracy. Nvidia says the model cuts the error rate by an average of 41 percent across eight test scenarios versus its predecessor Streaming Sortformer at a 1.04-second buffer. That number matters when the model has to run locally rather than in a data centre.
The same week brought a much larger project with a much narrower claim about automation. A team involving researchers from China's Fudan University studied its own work building Atria Dawn Preview, an agentic language model on a mixture-of-experts architecture with 744 billion parameters, according to The Decoder. The analysis covered more than 700 task logs from 56 participants. AI was used in 96.5 percent of tasks reviewed, and the median ratio of agent actions to human inputs rose from 11 to 28.5 over four weeks. The authors caution against reading that as autonomy: each human decision simply triggered more agent steps.
Participants were asked whether they could have completed their share of a task without AI at the same scope and quality. Of 455 completed AI-assisted tasks, 151 were rated infeasible without AI, roughly a third, spread across 27 of the 56 participants.
Humans still made 85.5 percent of decisions about methods and parameters, while AI made 9.2 percent, and humans made the final call on goals and scope in 93.4 percent of cases. When things went wrong across 588 tasks with a recorded difficulty, 76 percent moved forward through human intervention and 23 percent were solved by the agent alone. The paper warns that when every decision rests on a longer chain of agent work than any human can review, oversight degrades into rubber-stamping.
Cheap models, expensive questions
Cost is reshaping which models get used at all. CNBC reported on 26 September that Chinese models accounted for 57% to 67% of tokens used on OpenRouter in the week of 14 September, up from 6% to 13% in February, and reached 55% of tokens on Vercel in August from 11% in January. Peter Walker, head of insights at OpenRouter, told CNBC that Chinese open source models released this year can credibly handle advanced agentic tasks, especially coding, in a way that was not true in late 2025, and are inexpensive next to most models from US labs. Harpreet Arora, head of agentic infrastructure at Vercel, said price becomes compelling once a model clears the quality bar for a job.
Washington is watching. Two US House committees are investigating the adoption trend, and CNBC notes that US frontier models still attract more overall spending. Daniel Remler of the Center for a New American Security told the outlet the concern is that integration pulls countries into a Chinese technology sphere. The data CNBC cites is uneven: about half the tokens on OpenRouter come from US companies, while businesses in the 82 countries OpenRouter groups as the Global South send 67% of their tokens to Chinese models.
Small models sit awkwardly in this fight. They run on the device, which sidesteps some export-control arguments, but their reliability is exactly what the arXiv paper questions. Its routing result suggests the practical answer is not a better small model on its own, but a small model that knows when to ask for help.
Sources
6- 01Do small language models know what they don't know?EN
- 02Nvidia drops a free 100M-parameter model that identifies up to eight speakers in real timeEN
- 03AI agents do more of the work in model development, but humans still make the decisionsEN
- 04Chinese AI models surge in global popularity — and Washington is worriedEN
- 05OpenAI halts training of latest models as reports mount of AI agents going rogueEN
- 06OpenAI pauses training of its 'most capable models'EN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.