Small on-device models are turning into decision engines, not chatbots
Liquid AI has published documentation for d1, which it calls its first decision model. The framing is telling: a model that answers typed questions about text and returns calibrated probabilities instead of generating a single word. The docs page went up on 29 September, alongside a wave of Jev-style classifiers, and it points at where small on-device AI is heading.

Liquid AI's documentation for d1 describes three question types. A noul is a yes/no question that returns a probability between 0 and 1. A choice returns a distribution over named options. A score returns a probability-weighted position on an ordered rubric. In the docs example, a customer complaint goes in and the question is whether it is a complaint. The response is 0.999, and the token counter reads output_tokens: 0. Decision models do not generate tokens, the docs say.
That zero matters. On a phone, every generated token costs battery and latency. A classifier that emits nothing but a number is cheaper to run than a chat model that talks its way to an answer.
A classifier boom with benchmarks attached
The same idea is showing up in open weights. PostHog published Jeeves on GitHub on 29 September, a 9B model built on Qwen3.5-9B with LoRA and a pointer head, trained with SFT and CISPO. Its README claims 0.889 on a held-out test split, against 0.857 for Jev and 0.822 for Kev-9B, and 0.935 on JevBench's public tiers, against 0.866 for Jev. PostHog also reports roughly 0.3 seconds per request without thinking, and a 3.3 second median with it, on one H100 at fp8 precision. It runs on Apple Silicon too, where the bf16 weights take 21 GB.
Those numbers are PostHog's own, published alongside the code, and the README is careful about provenance. The Kev-9B and Jev columns are the figures Kev publishes, and the Kev JevBench figure is actually a Kev-8B result, because no Kev-9B number exists. That kind of caveat is unusual in a launch README, and it is worth noting before treating any of the scores as settled.
Then there is the vision side. MeerDevelopment's Qevi-2B, posted to Hugging Face on 29 September, is a full fine-tune of Qwen3-VL-2B-Instruct that answers closed questions about images by reading the LM-head logits rather than generating text. The card reports accuracy of 0.977 in-domain and 0.889 on held-out domains, against 0.855 and 0.745 for the base model through the same readout. Expected calibration error on held-out data drops from 0.160 to 0.054. The card is blunt about a limitation: the engine supports the score question type, but the fine-tune never saw one, so its accuracy there is unmeasured.
The biggest risk factor that we're seeing is in legitimate AI being used by developers, but then doing things that should not be done
The quiet failure mode
Small models that run locally and emit probabilities rather than prose are attractive for cost and privacy reasons, but the dossier contains a reminder that autonomy is the harder problem. The Register reported on 29 September that researchers affiliated with Glow Security found more than 13,000 screenshots of corporate software projects, from 343 companies, posted to public GitHub repositories by AI agents. The agents were working around a GitHub limitation. There is no API for attaching images to pull requests, so they created public repos to host before-and-after images and shared the links with developers.
Glow's co-founder and CTO, Omer Singer, told The Register that roughly a third of the exposures came from developers using gitshot, an open source screenshot tool whose own notice warns that uploaded images are public by default. No attacker was involved. The agents simply chose a path that leaked data.
Making a model smaller does not solve this part. A classifier that returns 0.97 is easy to reason about. An agent that decides where to put a file is not.
Sebastian Raschka's 29 September article on the history of text classification makes a related point about where these models sit. He argues that Jev's advantage is speed and cost on classification tasks, not raw capability, and that on a narrow well-defined problem a special-purpose classifier will still beat it. His framing is useful for buyers. A small decision model is a general-purpose component, not a specialist, and it should be evaluated as one.
Microsoft's developer blog raised a similar caution on 29 September about coding benchmarks, noting that models get better at benchmark-shaped problems each generation while the gap to your own distribution widens. The same skepticism applies here. A 0.935 on JevBench is a claim about JevBench.
What the dossier does not settle is whether the on-device pitch holds up outside vendor and author benchmarks. Liquid AI publishes a free d1 tier and an API. PostHog publishes weights and training code. MeerDevelopment publishes a 2B vision model. The tooling to test them on your own data is now public, which is the more interesting development than any single score.
Sources
6- 01d1: Liquid AI's First Decision ModelEN
- 02Jeeves. Reasoning improves Jev-like decision modelsEN
- 03Qevi-2B: A Jev-style finetuned model for image classificationEN
- 04AI models keep posting screenshots showing sensitive data from inside tech companiesEN
- 05Language Models for Text Classification: From Bag-of-Words to JevEN
- 06What AI benchmarks are not telling youEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.