Small models learn to answer without writing: decision models arrive on-device
On 29 September, Liquid AI published docs for d1, its first decision model. It returns calibrated probabilities in a single call and generates zero tokens. It is the newest entry in a fast-moving class of small, on-device classifiers that includes Jev, Qevi-2B and PostHog's Jeeves.

The pitch is narrow and specific: a model that does not write. Liquid AI's documentation, published on 29 September, describes how d1 evaluates a piece of state and returns probabilities across a fixed set of outcomes. The company's own example sends a customer complaint and gets back a single number, 0.999, for the question "Is this message a complaint?"
Liquid calls the format a decision model. It supports three question types, and the docs describe them in plain terms. A "noul" question is yes/no and returns a probability between 0 and 1. A "choice" question returns a distribution over named options. A "score" question returns a probability-weighted position on an ordered rubric. The API response includes an output_tokens field that is always zero.
Why the small models are suddenly interesting
Liquid is not alone in this corner. Sebastian Raschka's 29 September technical write-up on the history of text classification frames the recent Jev model as a classifier first and a cultural phenomenon second. His argument: Jev is not better or cheaper than a purpose-built classifier on a narrow task, but it is far more general than those task-specific models while staying much faster and cheaper than a frontier LLM.
That combination, mid-level generality at low cost per call, is what makes the on-device story plausible. A 2B or 9B model that returns a probability instead of a paragraph can run on hardware a frontier model cannot touch.
Qevi-2B, posted to Hugging Face on 29 September, is the image-side version. It is a full fine-tune of Qwen3-VL-2B-Instruct that answers closed questions about images by reading logits at the position where the model would start replying, not by generating text. The model card is blunt about the trade: "It is not a chat model. Ask it 'is there a ladder in this image?' and it returns P(Yes) = 0.97, not a sentence."
The card reports accuracy of 0.977 in-domain against 0.855 for the base model, and 0.889 on held-out domains against 0.745. Expected calibration error on held-out data drops from 0.160 to 0.054. The card also warns anyone tempted to stack temperature scaling on top: at T = 2.45 the held-out ECE returns to 0.160, wiping out the entire gain.
Speed is the other half of the claim. Qevi-2B says it is up to about 21x faster when asking many questions about one image. All three example questions share one encoding and are separated by a block-diagonal attention mask. The card says the answers are identical to asking them separately, verified bitwise.
"The biggest risk factor that we're seeing is in legitimate AI being used by developers, but then doing things that should not be done," said Omer Singer, co-founder and CTO of Glow Security, in an interview with The Register.
The catch: agents that leak while helping
That quote has little to do with classification accuracy, and everything to do with deployment. The Register reported on 29 September that researchers at Glow Security found more than 13,000 sensitive screenshots of corporate software projects, from 343 companies, posted to public GitHub repositories by AI agents. They call it PixelLeak.
The mechanism is mundane. GitHub has no API for uploading images to pull requests, so agents asked to show before-and-after interface shots put them in a public repo instead. About a third of the exposures came from developers using gitshot, an open source screenshot tool whose own privacy notice warns that uploaded images are public by default.
None of this is caused by small models. But it is the environment in which small models are being pitched: cheap enough to run constantly, embedded in developer tools, acting without a human in the loop. A classifier that runs on a laptop makes that loop tighter.
Reasoning, but only a little
PostHog's Jeeves, published on GitHub on 29 September, tries to fix the accuracy problem in Jev-like models without giving up the format. It is a 9B Jev-style classifier built on Qwen3.5-9B with LoRA and a pointer head, trained to reason before it decides, with a block-4 diffusion drafter.
The repo's numbers compare it against Kev-9B and Jev on data it was never trained on: 0.889 against 0.822 and 0.857. On JevBench's public tiers it reports 0.935 against 0.866 for Jev. The cost is latency. Without thinking, a request takes about 0.3 seconds. With thinking, the median is 3.3 seconds on one H100 at FP8 precision. The repo notes that truncating the chain shortens that.
The Jeeves table also contradicts the simple story that reasoning always helps. On MMLU-Pro it scores 0.739 against Jev's 0.840, and on MMLU 0.793 against 0.900. The authors do not average those away; they publish them side by side.
Microsoft's developer blog, in a 29 September post on Agent Experience, makes the general version of this point. Benchmark scores drive adoption, adoption drives optimization for the benchmark, and the score stops telling you about your own codebase. "A model that's excellent at Django might be mediocre at your internal framework that superficially resembles Django but works differently in critical ways," the post says. The same logic applies to a decision model tested on public tiers and then dropped into your ticket queue.
On-device is where these threads meet. Qevi-2B lists 21 GB of weights in bf16, and Jeeves says an Apple Silicon Mac needs 48 GB or more for the default caches, though FP8 cuts the weights to 11.5 GB and smaller cache settings fit a 48 GB machine. That is not a phone. It is a laptop, which is still a different deployment target from a data centre, and the reason the small-model label is doing real work here.
What none of these releases settles is whether the classification framing survives contact with messy production questions. A 0.97 is easy to threshold. A score question with unmeasured calibration, which is what Qevi-2B's card says about its own "score" type, is not.
Sources
6- 01d1: Liquid AI's First Decision ModelEN
- 02Qevi-2B: A Jev-style finetuned model for image classificationEN
- 03Jeeves. Reasoning improves Jev-like decision modelsEN
- 04AI models keep posting screenshots showing sensitive data from inside tech companiesEN
- 05Language Models for Text Classification: From Bag-of-Words to JevEN
- 06What AI benchmarks are not telling youEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.