Small language models move on-device, and the research is still catching up
Apple's OpenELM family spans 270 million to 3 billion parameters, and Google's Gemma 3 1B runs at 529MB. A 2026 paper finds token-level confidence signals near zero in 91% of dataset-model combinations.

Small language models run on a phone, not in a data centre. They stay under 3 billion parameters, and in roughly two years they have gone from a research curiosity to a shipping product. The pitch is simple: keep the model local, cut server costs, keep user data off someone else's hardware. The evidence is more mixed than the marketing suggests.
Apple is the clearest starting point. On 24 April 2024, Ars Technica reported that the company released eight models under the name OpenELM, short for Open-source Efficient Language Models, on Hugging Face under an Apple Sample Code License. Ars Technica noted the licence carries restrictions that may not fit the commonly accepted definition of open source, though the source code is available. The models range from 270 million to 3 billion parameters, in four pretrained and four instruction-tuned variants, with a 2048-token maximum context window. Apple's white paper says a "layer-wise scaling strategy" delivered a 2.36 percent accuracy improvement over Allen AI's OLMo 1B while using half as many pre-training tokens.
Ars Technica put that in context. Meta's Llama 3 family topped out at 70 billion parameters at the time, with a 400 billion version in development, and OpenAI's GPT-3 from 2020 shipped with 175 billion. Parameter count is a rough measure, not a verdict.
What the benchmark papers actually measure
SlimLM, posted to arXiv on 15 November 2024, took a narrower route: document assistance. The authors, Thang M. Pham and four co-authors, tested models from 125M to 7B parameters on a Samsung Galaxy S24. They pre-trained on SlimPajama-627B and fine-tuned on a dataset they built called DocAssist, for summarisation, question answering and suggestion tasks. Their abstract claims comparable or superior performance against existing small models, and it describes an Android application as a deployment reference. That is a useful data point, because it names the hardware, the task and the model sizes together. Most on-device claims do not.
Google's AI Edge team went wider. In a 20 May 2025 developer blog post, Mark Sherwood, Matthew Chan, Marissa Ikonomidis and Milen Ferev announced support for more than a dozen models, including Gemma 3 and Gemma 3n, hosted on a new LiteRT Hugging Face community. Gemma 3 1B is 529MB and can run up to 2,585 tokens per second pre-fill on the mobile GPU. Google says that lets it process up to a page of content in under a second. Gemma 3n, in 2B and 4B variants, is described as Gemma's first multimodal on-device small language model, supporting text, image, video and audio, with text and image available first. The post also introduces on-device RAG and function calling libraries, both available on Android at launch.
Then there is the sceptical layer. A paper titled "Do small language models know what they don't know?" was submitted to arXiv on 21 July 2026 by Prashant Mudgal. It evaluates seven approaches to entropy-based confidence across 7 model pairs and 5 NLU benchmarks. The headline finding is blunt. Token-level entropy is effectively blind in small models: it sits near zero whether the answer is correct or not, in 91% of dataset-model combinations. Semantic entropy, computed by sampling and clustering answers by meaning, recovers a usable signal. Routing uncertain queries to a larger expert model improved accuracy by up to 50 percentage points, with cross-family routing averaging +22.0% against +6.8% for same-family routing.
Reasoning without the running commentary
Science News reported on 22 September 2026 on BDH-CQ, a small experimental system submitted to arXiv on 10 August that solves some reasoning puzzles without writing out intermediate steps. Complexity scientist Zuzanna Stamirowska, CEO of AI company Pathway, told Science News: "Once the query arrives, the model doesn't write out its thinking in words at all." The model solved nearly three in 10 puzzles on the public ARC-AGI-1 evaluation set with two attempts. Stamirowska and colleagues estimate each puzzle query costs about $0.00070 to run, roughly one-eleventh of GPT-5.6 Luna on the same benchmark. Science News notes the two costs were calculated differently and the study has not been peer-reviewed.
Not everyone is sold. Computer scientist Yuntian Deng of the University of Waterloo told Science News the result is "an interesting efficiency result" but does not show the architecture is better than other approaches, and that without written reasoning the model is harder to inspect. Jonas Geiping of the ELLIS Institute Tübingen and the Max Planck Institute for Intelligent Systems called the approach "neat" but said BDH-CQ was built specifically for ARC-style problems, which makes direct comparison with general-purpose systems difficult.
Deployment is already outrunning the papers. Imbue's Bouncer, a Twitter feed filter published on 8 April 2026, uses an open-source model the company calls Gemma 4 26B-A4B for classification, served from Imbue's own data centre for now. Engineer Millan Philipose writes that the team is working to get it running directly on laptops and phones. That ambition, set against a 529MB model on a handset, is where most of the next round of claims will be tested.
Sources
6- 01Apple releases eight small AI language models aimed at on-device useEN
- 02SlimLM: An Efficient Small Language Model for On-Device Document AssistanceEN
- 03On-device small language models with multimodality, RAG, and Function CallingEN
- 04Do small language models know what they don't know?EN
- 05Can AI reason without words? A small model puts the idea to the testEN
- 06Show HN: Control your X/Twitter feed using a small on-device LLMEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.