Small language models move on-device: Apple, Google, Imbue and Cactus push local AI
Apple, Google and a crop of smaller vendors are pushing small language models onto phones, laptops and wearables, betting that inference can run locally instead of in a data centre.

Apple released eight source-available small language models in April 2024 under the name OpenELM. They range from 270 million to 3 billion parameters, according to Ars Technica. The models went up on Hugging Face under an Apple Sample Code License. Ars Technica noted that the licence may not meet the commonly accepted definition of open source because of its restrictions, even though the source code is available.
Four of the eight are pretrained and four are instruction-tuned. All eight share a 2048-token maximum context window. The training data mixes RefinedWeb, a de-duplicated version of PILE, a subset of RedPajama and a subset of Dolma v1.6. Apple says that totals around 1.8 trillion tokens.
Apple's white paper claims a layer-wise scaling strategy gave OpenELM a 2.36 percent accuracy improvement over Allen AI's OLMo 1B while using half as many pre-training tokens. The company also released CoreNet, the training library, plus reproducible training recipes so the weights can be replicated. Ars Technica described that as unusual for a major tech company at the time.
Parameter counts are a rough measure of capability, and Apple's models sit well below the largest systems. Meta's Llama 3 family had reached 70 billion parameters, with a 400 billion version in the works. OpenAI's GPT-3 shipped with 175 billion in 2020. The research bet is that small models can now match what much larger ones did a few years ago.
Google adds modalities, RAG and function calling
Google's AI Edge team said on 20 May 2025 that it had expanded on-device small language model support from four initial models to more than a dozen, including Gemma 3 and Gemma 3n, hosted on a new LiteRT Hugging Face community. Gemma 3n arrived as an early preview. Google described it as Gemma's first multimodal on-device small language model, with 2B and 4B parameter variants supporting text, image, video and audio inputs. Text and image were available on Hugging Face first, with audio to follow.
Google also put numbers on the compression work. It said int4 post-training quantisation can shrink language models by a factor of 2.5 to 4 times compared with bf16, while cutting latency and peak memory. Gemma 3 1B, at 529MB, can run up to 2,585 tokens per second pre-fill on a mobile GPU. Google says that is enough to process a page of content in under a second.
Two libraries shipped alongside the models. The AI Edge RAG library, available on Android at launch, lets a model draw on application-specific data without fine-tuning. Google says it can pick the most relevant pieces from 1,000 pages or 1,000 photos. The Function Calling library, also Android-first, lets a model decide when to call predefined functions. A sample app fills a medical form from dictated speech.
The reproducibility and transparency of large language models are crucial for advancing open research, ensuring the trustworthiness of results, and enabling investigations into data and model biases, as well as potential risks.
That line comes from Apple's OpenELM paper abstract, as quoted by Ars Technica. Apple also cautioned that because the models were trained on publicly sourced datasets, they may produce outputs that are inaccurate, harmful, biased or objectionable. Ars Technica reported that iOS 18, expected to be revealed in June at WWDC, was rumoured to include on-device AI features, with the possibility that Apple would hire Google or OpenAI for heavier off-device processing.
Benchmarks on a phone, and a filter for your feed
A November 2024 arXiv paper from researchers including Thang M. Pham presented SlimLM, a series of models tuned for document assistance on mobile. The authors ran experiments on a Samsung Galaxy S24 to find trade-offs between model size, from 125M to 7B parameters, context length and inference time. SlimLM was pretrained on SlimPajama-627B and fine-tuned on DocAssist, a dataset the authors built for summarisation, question answering and suggestion tasks. They also shipped an Android application.
Imbue's Bouncer takes a different angle. Published on 8 April 2026 and updated on 25 September 2026, it filters an X/Twitter feed from a plain-language description, using an open-source model the company calls Gemma 4 26B-A4B. Imbue says it is serving that model from its own data centre for now while working to run it directly on a laptop or phone. Bouncer is available as a Chrome extension, a Firefox add-on and an iPhone app. Reddit, LinkedIn and Safari support are listed as coming.
Cactus Compute's Needle 3, posted on 18 September 2026, goes smaller still. The company describes 29-121M parameter laddered models producing 9 to 29 MB CQ2-bit binaries, with speed of 400 to 4,000 tokens per second decode and 1 to 10,000 tokens per second prefill on a Raspberry Pi 5. Cactus says the 4-layer subnetwork can match DeepSeek V4 Flash when tuned on downstream tasks for one epoch.
Needle 3 targets tool calls, structured extraction and text embedding, with quoted use cases in smart homes, robots, phones, wearables, AR glasses, cars and PCs. Every response carries a confidence score from a calibrated head. The engine applies a floor of 0.1, below which calls are withheld into a suppressed_calls field.
One research result cuts against the enthusiasm. A paper submitted on 21 July 2026 by Prashant Mudgal found that token-level entropy is effectively blind in models under 3 billion parameters: in 91 percent of dataset-model combinations, mean token entropy was near zero regardless of whether the answer was correct. Semantic entropy, computed by sampling multiple answers and clustering them by meaning, recovered a usable signal. Routing uncertain queries to a larger expert model improved accuracy by up to 50 percentage points, with cross-family routing averaging +22.0 percent against +6.8 percent for same-family routing.
The paper's framing is blunt about what this means for local inference. The value of entropy methods in small models is not computational savings, it argues, but intelligent compute allocation: spending more tokens where they matter most. For now, that leaves the pitch for on-device AI resting on privacy, latency and cost, with the harder question of reliability still open.
Sources
6- 01Apple releases eight small AI language models aimed at on-device useEN
- 02SlimLM: An Efficient Small Language Model for On-Device Document AssistanceEN
- 03On-device small language models with multimodality, RAG, and Function CallingEN
- 04Do small language models know what they don't know?EN
- 05Show HN: Control your X/Twitter feed using a small on-device LLMEN
- 06Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 FlashEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.