Small language models move on-device, but confidence signals lag behind
Apple's OpenELM family spans 270 million to 3 billion parameters, and Google now ships Gemma 3n with text, image, video and audio inputs. A 2026 arXiv paper, though, reports that token-level entropy is effectively blind in models under 3 billion parameters.

On-device small language models have stopped being a research curiosity. In two years they have gone from proof-of-concept weights on Hugging Face to shipped libraries with retrieval, function calling and multimodal inputs. The pitch is consistent: keep inference on the phone or laptop, cut server costs, keep user data local.
The tooling that tells these models when they are wrong has not kept pace.
From eight tiny models to a library
Apple opened this phase on 24 April 2024, when it released eight source-available language models under the name OpenELM, short for Open-source Efficient Language Models, according to Ars Technica. They ranged from 270 million to 3 billion parameters, split into four pretrained and four instruction-tuned variants. Ars Technica put the size in context: Meta's Llama 3 family topped out at 70 billion parameters at the time, with a 400 billion version in the works, while OpenAI's GPT-3 from 2020 shipped with 175 billion.
Apple's claim was not just size. The company said its layer-wise scaling strategy allocated parameters more efficiently across layers, and its white paper reported a 2.36 percent accuracy improvement over Allen AI's OLMo 1B while using half as many pre-training tokens. Apple also published the CoreNet training library and reproducible training recipes, which Ars Technica called unusual for a major tech company. The models carried a 2048-token context window and were trained on RefinedWeb, a deduplicated PILE, a subset of RedPajama and a subset of Dolma v1.6, around 1.8 trillion tokens in total.
There exists the possibility of these models producing outputs that are inaccurate, harmful, biased, or objectionable in response to user prompts, Apple cautioned in its paper, as quoted by Ars Technica.
That caveat matters more as these models leave the lab.
Google turns SLMs into a platform
By May 2025 the framing had changed from releasing models to supporting an ecosystem. Google's AI Edge team said it had expanded on-device support from four initial models to more than a dozen, hosted on a new LiteRT Hugging Face community, including Gemma 3 and Gemma 3n.
Gemma 3n arrived as an early preview and was described as Gemma's first multimodal on-device small language model, with 2B and 4B parameter variants supporting text, image, video and audio. Google said text and image modalities were available on Hugging Face, with audio to follow. The company also pointed to Gemma 3 1B, at 529MB, which it said could run up to 2,585 tokens per second pre-fill on a mobile GPU, enough to process a page of content in under a second.
The more consequential part of the announcement was infrastructure. Google shipped an on-device Retrieval Augmented Generation library, available on Android at launch, and an on-device function calling library, also Android-first. The RAG library lets a model draw on application-specific data without fine-tuning, and Google said it could surface the most relevant pieces from 1000 pages of information or 1000 photos. The function calling library lets a model decide when to invoke predefined functions or APIs. Google's sample app fills a form from dictated speech: it converts voice to text, extracts fields and calls application functions.
Google also described new int4 post-training quantization schemes that it said can reduce model size by a factor of 2.5 to 4 times compared with bf16, while decreasing latency and peak memory consumption. That is the kind of engineering detail that decides whether a model ships or stays a demo.
SlimLM and the Samsung benchmark
Academic work has run in parallel. A paper submitted to arXiv on 15 November 2024, revised through 25 November, presented SlimLM, a series of small language models optimized for document assistance on mobile devices. The authors, Thang M. Pham, Phat T. Nguyen, Seunghyun Yoon, Viet Dac Lai, Franck Dernoncourt and Trung Bui, ran experiments on a Samsung Galaxy S24 and mapped trade-offs between model size, from 125M to 7B parameters, context length and inference time.
SlimLM was pre-trained on SlimPajama-627B and fine-tuned on DocAssist, a dataset the authors built for summarization, question answering and suggestion tasks. The abstract says the smallest model performed efficiently on the S24 while larger variants offered more capability within mobile constraints, and that the work compared favourably with existing small language models. The authors released an Android application as well, and framed the payoff as reduced server costs and better privacy through on-device processing.
The confidence problem
Then there is the question of whether these models know what they do not know. A paper submitted to arXiv on 21 July 2026 by Prashant Mudgal examined entropy-based confidence signals in small language models with fewer than 3 billion parameters, running entirely on consumer hardware. It evaluated seven approaches across 7 model pairs and 5 standard NLU benchmarks.
The headline finding is uncomfortable. Token-level entropy was effectively blind: in 91 percent of dataset-model combinations, mean token entropy was near zero regardless of whether the answer was correct. The paper says that renders token-based confidence signals unusable at this scale.
Semantic entropy fared better. By generating multiple samples, clustering answers by meaning and measuring distributional uncertainty, the method recovered a usable confidence signal. Using it to route uncertain queries to a larger expert model produced accuracy improvements of up to 50 percentage points. Cross-family routing, the paper gives the example of SmolLM 360M to Phi-3.5-mini, averaged a 22.0 percent improvement against 6.8 percent for same-family routing. The author's conclusion is that the value of entropy methods in small models is not computational savings but intelligent compute allocation, spending more tokens where they matter most.
Reasoning without words
A separate line of work asks whether small models need to write out their reasoning at all. Science News reported on 22 September 2026 that a small experimental system called BDH-CQ, where DH stands for Dragon Hatchling, can solve some reasoning puzzles without spelling out intermediate steps. The paper was submitted to arXiv on 10 August and has not been peer-reviewed.
Instead of keeping examples in context, BDH-CQ uses each example to update a fixed-size memory. Zuzanna Stamirowska, a complexity scientist and CEO of the AI company Pathway, told Science News that once the query arrives the model does not write out its thinking in words at all. "Nothing in between ever converts into language," she said.
Stamirowska and colleagues reported that the model solved nearly three in 10 puzzles on the public ARC-AGI-1 evaluation set when given two attempts. It handled some shape-turning and shape-moving tasks but struggled with colour changes and some combinations of rules, and did better on harder ordering and nesting puzzles after seeing an example of similar difficulty.
The cost claim is striking. The researchers estimate each puzzle query costs about $0.00070 to run, roughly one-eleventh as much as GPT-5.6 Luna on the same benchmark, though Science News notes the two costs were calculated differently. Outside researchers were measured in their reaction. Yuntian Deng of the University of Waterloo called it an interesting efficiency result but said it does not show the underlying architecture is better than other approaches, and noted that without written-out reasoning the model is harder to inspect. Jonas Geiping of the ELLIS Institute Tübingen and the Max Planck Institute for Intelligent Systems said the model was built specifically for ARC-style problems, making direct comparison with general-purpose systems difficult, though he called the approach neat and noted it can tackle test problems without retraining.
Shipping ahead of understanding
Imbue's Bouncer, a Twitter feed filter published on 8 April 2026, shows what deployment looks like in practice. The tool classifies posts using an open-source model the blog post calls Gemma 4 26B-A4B, which it says is large enough to understand context yet small enough to keep up with scrolling speed. Imbue says it is currently serving the model from its own data center while working to run it directly on laptops and phones, and that Bouncer is available as a Chrome extension, a Firefox add-on and an iPhone app, with Reddit, LinkedIn and Safari support promised.
The pattern across these projects is consistent. Capability and distribution are moving faster than evaluation. Apple shipped weights with a candid note about inaccurate or biased outputs. Google shipped retrieval and function calling so models can act on application data. SlimLM benchmarked document tasks on real hardware. Mudgal's paper found that the cheapest confidence signal does not work below 3 billion parameters, and that a more expensive one does. BDH-CQ suggests some reasoning can happen without language, while researchers caution that unspoken reasoning is harder to audit.
Sources
6- 01Apple releases eight small AI language models aimed at on-device useEN
- 02SlimLM: An Efficient Small Language Model for On-Device Document AssistanceEN
- 03On-device small language models with multimodality, RAG, and Function CallingEN
- 04Do small language models know what they don't know?EN
- 05Can AI reason without words? A small model puts the idea to the testEN
- 06Show HN: Control your X/Twitter feed using a small on-device LLMEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.