On-device small language models move from demos to deployed tools
Apple's OpenELM release in April 2024 put eight small language models between 270 million and 3 billion parameters on Hugging Face. Two years later, the on-device push has widened to multimodality, retrieval and function calling.

The pitch behind small language models is simple: keep the inference on the phone, keep the data off the server. Apple made that argument concrete on 24 April 2024, when Ars Technica reported the release of OpenELM, eight source-available models ranging from 270 million to 3 billion parameters.
Apple shipped them in two flavours, four pretrained and four instruction-tuned, under an Apple Sample Code License. Ars Technica noted that the licence carries restrictions, so OpenELM may not meet the commonly accepted definition of open source even though the source code is published. The models use a 2048-token maximum context window. They were trained on RefinedWeb, a de-duplicated PILE, and subsets of RedPajama and Dolma v1.6, which Apple says totals around 1.8 trillion tokens.
The scale comparison matters. Meta's Llama 3 family topped out at 70 billion parameters at the time, and OpenAI's GPT-3 from 2020 shipped with 175 billion. Apple's largest OpenELM is 3 billion.
"The reproducibility and transparency of large language models are crucial for advancing open research, ensuring the trustworthiness of results, and enabling investigations into data and model biases, as well as potential risks."
That line comes from Apple's OpenELM paper abstract, quoted by Ars Technica. Apple also released CoreNet, the training library, plus reproducible training recipes that let the weights be replicated, which Ars Technica described as unusual for a major tech company at that point. The same reporting flags Apple's own caveat: because the training data is publicly sourced, the models may produce inaccurate, harmful, biased or objectionable output.
Research followed the hardware. A paper submitted to arXiv on 15 November 2024 presented SlimLM, a series of models from 125M to 7B parameters tuned for document assistance on mobile devices. The authors ran experiments on a Samsung Galaxy S24 to find trade-offs between model size, context length and inference time, pretrained on SlimPajama-627B and fine-tuned on a dataset they built called DocAssist for summarisation, question answering and suggestion. They released an Android application alongside the paper and framed the goal as reducing server costs and improving privacy through on-device processing.
Google adds retrieval and function calling
Google's AI Edge team moved the category further in a post published on 20 May 2025. The company said it was expanding support from four initial on-device small language models to more than a dozen, including Gemma 3 and an early preview of Gemma 3n, hosted on a LiteRT Hugging Face community.
Gemma 3n is described as Gemma's first multimodal on-device small language model, supporting text, image, video and audio inputs across 2B and 4B parameter variants. Text and image modalities were available on Hugging Face, with audio to follow. Google also claimed that int4 post-training quantisation can shrink language models by a factor of 2.5 to 4 times compared with bf16, cutting latency and peak memory. Its Gemma 3 1B model, at 529MB, runs up to 2,585 tokens per second pre-fill on the mobile GPU, enough to process a page of content in under a second, according to the post.
The more consequential additions were libraries rather than weights. Google said its RAG library can pull the most relevant few pieces of data from 1,000 pages of information or 1,000 photos to ground a model without fine-tuning, and that its function calling library lets a model decide when to invoke predefined functions inside an app. Both were available on Android at launch, with other platforms to follow.
Confidence remains the open question. A paper submitted to arXiv on 21 July 2026 tested whether entropy-based signals can tell when a small model is wrong. Across seven model pairs and five NLU benchmarks, the author found token-level entropy effectively blind in models under 3 billion parameters: in 91% of dataset-model combinations, mean token entropy sat near zero regardless of whether the answer was correct. Semantic entropy recovered a usable signal, and routing uncertain queries to a larger expert model improved accuracy by up to 50 percentage points. Cross-family routing averaged a 22.0% gain against 6.8% for same-family routing, suggesting expert model quality matters more than architectural compatibility.
Commercial tools are already shipping. Cactus's Needle 3, posted on 18 September 2026, describes a 29-121M parameter model family with 9 to 29 MB CQ2-bit binaries, tool calling, structured extraction and text embedding, and claims its 4-layer subnetwork can match DeepSeek V4 Flash when tuned on downstream tasks for one epoch. The company reports 400-4k tokens per second decode on a Raspberry Pi 5.
Apple, for its part, has not integrated OpenELM into consumer devices. Ars Technica reported that iOS 18, expected to be revealed in June at WWDC, was rumoured to include on-device AI features, with the possibility that Apple would hire Google or OpenAI for heavier off-device processing. Two years on, the small models are real. The product decisions are still catching up.
Sources
5- 01Apple releases eight small AI language models aimed at on-device useEN
- 02SlimLM: An Efficient Small Language Model for On-Device Document AssistanceEN
- 03On-device small language models with multimodality, RAG, and Function CallingEN
- 04Do small language models know what they don't know?EN
- 05Needle 3 - 8-29 MB foundation model for tiny devicesEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.