Decision models go small: Qevi-2B, Liquid d1 and Jeeves push classification on-device
A 2B-parameter vision model that answers yes/no questions by reading its own logits, not by generating text, was published to Hugging Face on 29 September, claiming 0.977 in-domain accuracy and up to 21x faster inference when many questions are asked of one image.

The model is called Qevi-2B, a full fine-tune of Qwen3-VL-2B-Instruct. Its authors describe it as "not a chat model": ask whether a ladder is in a photo and it returns P(Yes) = 0.97 rather than a sentence. The published card puts in-domain accuracy at 0.977 against 0.855 for the base model, held-out domains at 0.889 against 0.745, and expected calibration error on held-out data at 0.054 versus 0.160. The speed claim rests on question packing. Multiple questions share a single encoding of the image, isolated by a block-diagonal attention mask, verified bitwise against asking them separately.
That is one artifact in a fortnight-old category that now has SDKs, benchmarks, a rival benchmark suite and at least one open-source clone. TypeSafe shipped Jev on 15 September, according to Casco, which timed the model's rise to meme status at "within forty-eight hours."
Liquid puts a name on the primitives
Liquid AI published documentation for d1, which it calls its first decision model, on 29 September. The docs are unusually explicit about what these systems are for. Instead of generating tokens one at a time, a decision model evaluates a state and returns calibrated probabilities over a fixed set of outcomes "in a single call with zero generated tokens."
Liquid defines three question types that have become the de facto interface. Noul is a yes/no question returning a probability between 0 and 1; the docs use a spam check returning 0.92 as the example. Choice returns a distribution over named options, such as {"billing": 0.65, "technical": 0.30, "account": 0.05} for ticket routing. Score returns a probability-weighted position on an ordered rubric, so "how urgent is this?" across three levels can come back as 1.85, between Medium and High. Liquid's guidance is to pick the type by what your code does next: branch on a category, use Choice; compare against a threshold, use Score; gate on a boolean, use Noul.
The company also warns against a common conflation. A noul at 0.5, per the documentation, means maximum uncertainty between yes and no and "says nothing about degree or intensity." Measuring degree requires Score with defined levels.
Jeeves: reasoning on top, and a 3.3 second bill
The most interesting pushback on the category came from PostHog, which published Jeeves on 29 September: a 9B Jev-like model built on Qwen3.5-9B with LoRA and a pointer head, trained with SFT and CISPO so that it reasons before it decides. The repository is blunt about the tradeoff it is correcting. "Jev-like models give calibrated decision probabilities, but at low accuracy," the README says, which is why many pipelines already fall back to a reasoning model.
Jeeves reports 0.889 on a held-out and out-of-domain test split of 2,962 items, against 0.857 for Jev and 0.822 for Kev-9B, with the caveat that the Kev and Jev columns are numbers Kev publishes. On JevBench's 231 public items it claims 0.935 against 0.866 for Jev; on the 111-item hard tier, 0.865 against 0.730. Its calibration error on public JevBench items is 0.037. The same checkpoint scores 0.804 without thinking against 0.840 with it.
Then the cost. PostHog measures about 0.3 seconds per request without thinking and a 3.3 second median with it, on one H100 at FP8 precision, and notes that truncating the chain shortens that. Inference also runs on Apple Silicon with 48GB or more, which matters if the point of a small decision model is to keep the work local.
The base model numbers in the Jeeves table are not all like-for-like: the repository marks its Kev-9B JevBench figure with an asterisk noting that no Kev-9B JevBench result is published and that the numbers shown are Kev-8B on Qwen3.
Who actually pays for a classifier
The clearest production argument comes from Evan Schwartz, who runs Scour, a personalized content feed. In a 29 September post he describes asking 54 questions of roughly 1.1 million documents per month: 50 nouls, 3 choices and 1 score, about 2,150 tokens including roughly 260 tokens of fixed overhead. His questions are now about 88% of the input tokens, sent verbatim with every request. His total spend is under $150 per month, and he says running the same content through even the cheapest LLM would be prohibitive for a bootstrapped project.
His ask is specific: prompt caching, or the ability to register reusable question sets.
He notes he does not batch multiple posts per request because of the documented warning that accuracy falls as the state grows with unrelated content, and suggests an API that takes many questions and many inputs as an alternative that avoids the packing penalty. Others are already filling the gaps around TypeSafe. A GitHub project called Jeb, updated on 29 September, turns any OpenAI-compatible endpoint into a decision model by reading token logprobs and returning structured JSON at POST /v1/systemone, with the same request format over CLI. Its README is candid about the prerequisite: an endpoint can be OpenAI-compatible for ordinary chat while lacking the log probabilities Jeb needs, so check that capability before picking a provider.
The benchmark problem arrives early
Casco, a security firm, ran 2,449 findings through eight models including Jev, Claude and GPT, comparing CVSS 3.1 severity scores against its own published references. Every model overestimated on average. Jev assigned 550 critical scores to a dataset containing 55 published critical findings, with 500 of those 550 below critical in Casco's reference. GPT-6 Astra was the most accurate overall at 1.92 points mean absolute error; Jev ranked seventh at 3.40, ahead of Haiku at 3.94.
The signed errors all point the same way: Jev +3.30 points, Haiku +3.89, Astra +1.18, on a ten-point scale. Casco's framing is that structured, calibrated output did not make severity judgments more accurate against its reference, and that the resulting "security slop" costs engineers time. The study used the same finding evidence for every model, without Casco's scoring policy or internal application context, so it measures the judgment call rather than the arithmetic.
There is also a maintenance question nobody has answered. Jeeves is a fine-tune on Qwen3.5-9B with a diffusion drafter and a Jev-compatible API; Qevi-2B is a fine-tune of Qwen3-VL-2B. Both depend on base models that will be superseded. The category's pitch is that you can add questions without retraining, which is true of the API surface but not of the weights.
On 29 September, Sebastian Raschka published a long technical history of text classification from bag-of-words through RNNs, CNNs and transformers to what he calls Jev-like APIs, explicitly to "demystify the hype." His conclusion is the least exciting and probably the most useful one: for a narrow, well-defined problem, a Jev-style model will not beat a special-purpose classifier on accuracy, speed or cost. Its selling point is generality within the classification slot, and that is a much smaller claim than the reaction suggests.
Sources
12- 01Qevi-2B: A Jev-style finetuned model for image classificationEN
- 02d1: Liquid AI's First Decision ModelEN
- 03Jeeves. Reasoning improves Jev-like decision modelsEN
- 04Please add prompt caching to Jev-style modelsEN
- 05Jeb: Turn any OpenAI API into a decision modelEN
- 06Every model (incl. Jev) we tested inflates security finding severityEN
- 07Language Models for Text Classification: From Bag-of-Words to JevEN
- 08The System One models ecosystemEN
- 09AI models keep posting screenshots showing sensitive data from inside companiesEN
- 10AI Models Fail at Physics: Coding Harnesses Are to BlameEN
- 11How many tasks does it take to trust a cheaper model?EN
- 12Gates Foundation 5 Year Goal: Help 3B People to Use AI in Their LanguageEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.