Decision models, not small language models, are quietly taking over on-device AI
Liquid AI published documentation for d1, its first decision model family, on 29 September. The docs describe a classifier that returns calibrated probabilities with zero generated tokens. The same day, PostHog released Jeeves, a 9B reasoning classifier, and a 2B vision fine-tune called Qevi-2B appeared on Hugging Face.

The pattern in this week's releases is not smaller chatbots. It is models that answer typed, closed questions and return numbers, not sentences.
Liquid AI's documentation, posted on 29 September, lays out the pitch directly. A decision model, it says, "evaluates a situation and returns calibrated probabilities across a fixed set of outcomes in a single call with zero generated tokens." The API supports three question types. A Noul is a yes/no question returning a probability between 0 and 1. A Choice returns a distribution over named options. A Score returns a probability-weighted position on an ordered rubric. In the sample response in Liquid's docs, a customer complaint question returns a Noul of 0.999 and an output token count of 0.
That last number is the whole point. On-device inference is usually a story about shrinking a generative model until it fits on a phone. Decision models skip the generation step entirely, which removes the dominant cost of running a model locally.
Three releases, one shape
PostHog's jeeves repository, last updated on 29 September, describes "a reasoning Jev-style classifier with a diffusion drafter, trained with SFT and CISPO." It is a 9B Qwen3.5-9B fine-tune with LoRA and a pointer head that, per the repo, "thinks before it decides." The published numbers claim 0.889 on a held-out test split against 0.857 for Jev and 0.822 for Kev-9B, and 0.935 on JevBench's public tiers against 0.866 for Jev. PostHog reports about 0.3 seconds per request without thinking and a 3.3 second median with it, on one H100 at fp8 precision.
Also on 29 September, a 2B vision model called Qevi-2B was published on Hugging Face by MeerDevelopment. It is a full fine-tune of Qwen3-VL-2B-Instruct that "answers typed, closed questions about images by reading the model's own logits instead of generating text." The model card is blunt about the trade. Ask it whether there is a ladder in an image and it returns P(Yes) = 0.97, not a sentence. In exchange it reports 0.889 accuracy on held-out domains against 0.745 for the base model, and an expected calibration error of 0.054 against 0.160.
Qevi-2B's card also warns that the model must be used through a logit readout, not .generate(). Calling generate and parsing the text will not reproduce the published figures, because that is not the path the model was fine-tuned on. It is a small detail with a large implication: these systems are not interchangeable with the chat models developers already know how to call.
Why the calibration claim matters
Sebastian Raschka's 29 September write-up on the history of text classification puts the current wave in context. He traces the line from bag-of-words through naive Bayes, logistic regression, RNNs and transformers, and notes that a classifier's advantage is speed and cost rather than generality. His framing of Jev is careful: "Jev is essentially a text classifier," but "Jev is also not 'just' a text classifier."
The distinction he draws is between narrow task-specific models and something general enough to handle many classification jobs without retraining. Qevi-2B and Jeeves both sit in that middle ground. Neither is a general assistant. Both claim to beat purpose-built baselines on tasks they were not trained on.
The calibration numbers are the part worth scrutinising. Qevi-2B's card reports that temperature scaling makes its calibration worse, not better, because label smoothing during fine-tuning already removed the base model's overconfidence. At T = 2.45 the held-out ECE returns to 0.160, exactly the base model's figure. The card also states plainly that earlier versions recommended 2.45, 1.9 and 3.0, and that those values were fitted against the base model before fine-tuning and should not be used.
That is an unusually candid correction to publish on a model card. It also undercuts a common assumption: that a small model's probabilities need post-hoc rescaling to be usable as thresholds. Here the raw softmax is the calibrated output.
Jeeves takes the opposite bet. Its README states that Jev-like models "give calibrated decision probabilities, but at low accuracy," and that many pipelines therefore fall back to a reasoning model. Jeeves trains reasoning into the classifier itself. The trade is latency: 0.3 seconds without thinking, 3.3 seconds with it, on an H100. Without thinking the same checkpoint scores 0.804 on PostHog's 2,962-item test split, against 0.840 with it.
The on-device angle is memory, not parameter count
Qevi-2B's practical claim is throughput: up to roughly 21 times faster inference when asking many questions about a single image, because all questions share one encoding and are isolated by a block-diagonal attention mask. The card says the answers are identical to asking them separately, verified bitwise.
Jeeves is less forgiving. It needs Python 3.12 and a CUDA GPU, though inference also runs on Apple Silicon. On a Mac the weights use 21 GB in bf16 and the default caches use another 28 GB, so a 48 GB machine starts to swap. PostHog suggests cutting the caches to --max-rows 4 --max-len 4096, or using the fp8 weights at 11.5 GB. On an M4 Pro, one question thinks at about 20 tokens per second, rising to about 38 in fp8.
So the smallest of these three is the only one that plausibly runs on a phone today. The others are server-side classifiers that happen to be cheaper than calling a frontier model.
Liquid's d1 documentation describes the standard deployment path: an API key prefixed with liquid_, a POST to a decisions endpoint, and SDKs for Python and JavaScript. The free tier is called d1:free. There is no on-device binary described in the documentation.
The benchmark caveat
Microsoft's developer blog weighed in on 29 September with an argument that applies directly here. The post, part of a series on making AI coding agents work with proprietary technology, cites Goodhart's law. Benchmark scores drive adoption, it points out, adoption creates pressure to optimise for the benchmarks, and the benchmark therefore becomes less representative with each cycle.
Its specific claim is that a model scoring 92% on SWE-bench is demonstrably good at resolving well-documented issues in popular repositories, and that this says nothing about whether it will produce correct code against an internal library. "Benchmarks sample from a distribution," the post says. "Your work lives in a different one."
Substitute JevBench or MMLU-Pro for SWE-bench and the warning holds. PostHog publishes Jeeves' per-benchmark breakdown, and it is not uniformly better than Jev. Jeeves loses on MMLU (0.793 against 0.900), on MMLU-Pro 10-way (0.739 against 0.840) and on QNLI (0.913 against 0.925), while winning on PAWS, SciQ, Emotion and buried state. It also reports 0.055 on unknowable questions answered at p ≥ 0.9, against 0.090 for Jev, where lower is better.
Those are not numbers a marketing page would lead with. They are, however, the numbers a team evaluating a classifier for a specific pipeline needs.
What is actually new
None of the underlying techniques are novel. Naive Bayes classifiers date back to at least 1961, according to Raschka, where they were used to sort computer abstracts. Bag-of-words plus logistic regression powered spam filters for decades. What is new is that the same interface, a typed question returning a calibrated probability, now comes attached to models that generalise across domains without task-specific training.
Whether that generalisation survives contact with a real deployment is the open question. Glow Security's 29 September disclosure is a useful reminder that AI systems deployed by developers do unexpected things, though it concerns coding agents rather than classifiers. The firm found more than 13,000 sensitive screenshots from 343 organisations posted to public GitHub repos by AI agents working around an upload limitation.
For teams weighing a small on-device classifier, the practical checklist from this week's releases is short. Check whether the model is meant to be called through a logit readout or through generate. Check whether the published calibration figures were measured at the temperature you intend to use. And check the per-benchmark table, not the headline.
Sources
6- 01d1: Liquid AI's First Decision ModelEN
- 02Jeeves. Reasoning improves Jev-like decision modelsEN
- 03Qevi-2B: A Jev-style finetuned model for image classificationEN
- 04Language Models for Text Classification: From Bag-of-Words to JevEN
- 05What AI benchmarks are not telling youEN
- 06AI models keep posting screenshots showing sensitive data from inside tech companiesEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.