Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Small language models move on device: what the numbers actually show

Small language models are moving out of the data centre and onto phones, but the published results point to uneven progress. Apple's OpenELM line spans 270 million to 3 billion parameters, while one Google model is a 529MB download that runs up to 2,585 tokens per second on a mobile GPU.

AI & modelsAnalysisGrace OkonkwoPublished: 27 September 20264 min readSources 5
Small language models move on device: what the numbers actually show

Apple released eight small language models on 24 April 2024 under the collective name OpenELM, short for Open-source Efficient Language Models, according to Ars Technica. They range from 270 million to 3 billion parameters and sit on Hugging Face under an Apple Sample Code License. Ars Technica noted that the licence carries restrictions, so the release may not meet the usual definition of open source even though the source code is available.

The eight models come in two flavours: four pretrained and four instruction-tuned. The maximum context window is 2,048 tokens. Apple says the training data came from RefinedWeb, a version of PILE with duplications removed, a subset of RedPajama and a subset of Dolma v1.6, totalling around 1.8 trillion tokens. The company's white paper says its layer-wise scaling strategy produced a 2.36 percent accuracy improvement over Allen AI's OLMo 1B while using half as many pre-training tokens. Apple also published CoreNet, the training library, plus reproducible training recipes. The paper abstract states that reproducibility and transparency of large language models are essential for advancing open research.

Scale is the obvious contrast. Ars Technica put the largest Meta Llama 3 release at 70 billion parameters, with a 400 billion version on the way, and OpenAI's 2020 GPT-3 at 175 billion. Parameter count is a rough measure of capability, not a verdict.

What phone benchmarks show

SlimLM, submitted to arXiv on 15 November 2024 and revised on 25 November, is a series of small language models for document assistance. The authors ran experiments on a Samsung Galaxy S24 and looked at the trade-offs between model size (125M to 7B parameters), context length and inference time. SlimLM was pretrained on SlimPajama-627B and fine-tuned on DocAssist, a dataset the authors built for summarisation, question answering and suggestion tasks. An Android application accompanies the paper. The authors say they want lower server costs and better privacy through on-device processing.

Google's contribution, described on its Developers Blog on 20 May 2025, is more product-shaped. Google AI Edge expanded support from four initial models to over a dozen, including Gemma 3 and Gemma 3n, hosted on a LiteRT Hugging Face community. Gemma 3 1B is 529MB and can run up to 2,585 tokens per second pre-fill on a mobile GPU. Google says that is enough to process a page of content in under a second. Gemma 3n, in an early preview, is Gemma's first multimodal on-device small language model, with 2B and 4B parameter variants supporting text, image, video and audio inputs. Text and image are on Hugging Face, with audio to follow.

Google also shipped an on-device RAG library, available on Android, and an on-device function calling library, also Android first. The company claims int4 post-training quantisation can shrink language models by a factor of 2.5 to 4X compared with bf16, while cutting latency and peak memory.

"Once the query arrives, the model doesn't write out its thinking in words at all. Nothing in between ever converts into language," complexity scientist Zuzanna Stamirowska, CEO of AI company Pathway, told Science News.

That quote concerns BDH-CQ, an experimental system described in a paper submitted on 10 August to arXiv and reported by Science News on 22 September. The model updates a fixed-size memory with each example instead of keeping examples in context. Stamirowska and colleagues say it solved nearly three in 10 puzzles on the public ARC-AGI-1 evaluation set with two attempts. They estimate each query costs about $0.00070, roughly one-eleventh of GPT-5.6 Luna on the same benchmark. Science News notes the two costs were calculated differently and the study has not been peer-reviewed.

Outside research labs, the claims get harder to verify. Bouncer, a Twitter feed filter published by Imbue on 8 April 2026, uses an open-source model the blog calls Gemma 4 26B-A4B for classification. The post says the model is currently served from Imbue's own data centre, with work under way to run it on laptops and phones.

Scepticism is warranted, and some of it comes from researchers themselves. Yuntian Deng of the University of Waterloo told Science News that BDH-CQ is "an interesting efficiency result" but does not show its architecture is better than other approaches. He said more testing is needed to separate the design from the training method, and that a model which does not write out its reasoning is harder to inspect. Jonas Geiping of the ELLIS Institute Tübingen and the Max Planck Institute for Intelligent Systems said BDH-CQ is built specifically for ARC-style problems, making direct comparison with general-purpose systems difficult. He still called the approach "neat".

The unresolved question is how much of this survives contact with a shipping product. Apple, by its own framing in the OpenELM paper, treats the models as research releases. Google's own caveat is that Gemma 3n is an early preview. The gap between a benchmark table and a phone in a pocket remains the part nobody has published numbers for.

Comments 0

Sources

5
  1. 01Apple releases eight small AI language models aimed at on-device useEN
  2. 02SlimLM: An Efficient Small Language Model for On-Device Document AssistanceEN
  3. 03On-device small language models with multimodality, RAG, and Function CallingEN
  4. 04Can AI reason without words? A small model puts the idea to the testEN
  5. 05Show HN: Control your X/Twitter feed using a small on-device LLMEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.