Jev-style decision models move on-device as Qevi-2B and d1 target closed questions
A 2-billion-parameter image classifier posted to Hugging Face on Tuesday claims it answers closed questions about photos up to 21 times faster than the base model it was fine-tuned from. It is the latest sign that the Jev-style decision-model format is spreading beyond text.

The Qevi-2B model card, published on 29 September, describes a full fine-tune of Qwen3-VL-2B-Instruct that never generates a sentence. Ask it whether a ladder is in an image and it returns P(Yes) = 0.97, read directly from the language-model head logits at the position where the base model would have started its reply. The card states that calling .generate() and parsing the text will not reproduce the reported numbers, because that is not the path the model was fine-tuned on.
The claimed gains are in accuracy and calibration rather than raw capability.
Against the base Qwen3-VL-2B-Instruct run through the identical readout path on the same images, Qevi-2B scores 0.977 in-domain versus 0.855, and 0.889 on 12 held-out image domains versus 0.745. Expected calibration error on held-out data falls from 0.160 to 0.054.
One forward pass, many questions
The speed claim comes from question packing. The bundled qevi engine shares a single encoding of the image across several questions and isolates them with a block-diagonal attention mask. The card says the answers are identical to asking the questions separately, verified bitwise. The repository also warns that the score question type, an ordered rating on a rubric, was never in the training corpus, which contained 57,000 noul (yes/no) and 28,500 choice questions. Score accuracy and calibration are described as unmeasured and untested.
That vocabulary of question types is not Qevi's invention. Liquid AI's decision-model documentation, updated on 29 September, defines the same three shapes: noul for a yes/no probability between 0 and 1, choice for a distribution over named options, and score for a probability-weighted position on an ordered scale. The docs are explicit that decision models generate zero tokens. A sample complaint-classification request returns noul: 0.999 with output_tokens: 0.
Liquid's own framing of when to use which type is worth quoting, because it maps onto the on-device pitch. "If your code branches on a category, use a Choice. If it compares against a threshold on a continuum, use a Score. If it gates on a boolean, use a Noul," the documentation says. A classifier that returns a calibrated float is easier to threshold in a mobile app than a generative model whose answer has to be parsed.
The reasoning version runs the other way
Not everyone is pushing toward fewer parameters and less deliberation. PostHog's Jeeves repository, published on GitHub on 29 September, takes a 9B Jev-like model (Qwen3.5-9B with LoRA and a pointer head) and trains it with CISPO to think before it decides. The README reports 0.889 on a held-out test split versus 0.857 for Jev and 0.822 for Kev-9B, and 0.935 on JevBench's public tiers versus 0.866 for Jev.
The trade-off is latency.
Jeeves takes about 0.3 seconds per request without thinking and a 3.3 second median with it on one H100 at FP8 precision. The README notes the same checkpoint scores 0.804 on the authors' test split without thinking against 0.840 with it. On an M4 Pro, one question thinks at roughly 20 tokens per second, or about 38 in FP8. Weights use 21 GB in bf16, or 11.5 GB quantised, which the README says needs a 48 GB Mac and smaller caches to avoid swapping.
That is a different deployment target from Qevi-2B, and the two projects are pulling in opposite directions on the same underlying format.
Where the format came from
Sebastian Raschka's 29 September explainer traces the lineage from bag-of-words representations and naive Bayes through recurrent and convolutional networks to transformer-era classifiers, and argues Jev sits in an awkward middle. State-of-the-art LLMs can do the same classification tasks "while also being capable of much more general decision-making", he writes, but Jev handles them faster and more cheaply. At the narrow end, a purpose-built classifier will match it on a well-defined problem. Raschka states he is not affiliated with Jev and was not offered free access.
His history also supplies a useful sanity check on the hype. Text classification with naive Bayes goes back to at least 1961, where it was used to sort computer abstracts, and Raschka notes that Gmail's original spam filter allegedly used a naive Bayes model over a bag-of-words representation. Reading logits at a fixed position is a newer plumbing choice, not a new idea about what classification is.
Calibration is the claim to watch
For on-device use, the interesting number in the Qevi-2B card is not accuracy but the calibration table. The authors report that applying temperature scaling on top of the fine-tune makes calibration worse, because label smoothing during training already removed the base model's overconfidence: ECE is 0.054 at T = 1.0, 0.051 at T = 1.5, and back to 0.160 at T = 2.45, which is essentially the base model's held-out figure. Earlier versions of the card recommending 2.45 were fitted against the base model before fine-tuning, and the card now says they should not be used.
Two cautions apply.
Both Qevi-2B and Jeeves publish their own numbers, and the Jev and Kev columns in the Jeeves README are described as the numbers Kev publishes, restricted to the same public benchmark items; the sealed judge tier is excluded. Independent replication of any of these figures is not in the dossier.
There is also a live argument about what benchmark scores mean at all. A Microsoft for Developers post from 29 September quotes Goodhart's law, "when a measure becomes a target, it ceases to be a good measure", and points at data overlap between public benchmark tasks and training corpora. A model scoring 92% on SWE-bench, the post argues, "says however nothing about whether that same model will produce correct code when working with your internal auth library". The same scepticism transfers to JevBench tiers.
The direction of travel, though, is clear enough. A 2B vision model returning thresholded probabilities from logits, a 9B text model that reasons before it classifies, and a hosted API that bills zero output tokens are three implementations of the same bet: that many production decisions do not need generated prose at all.
Sources
6- 01Qevi-2B: A Jev-style finetuned model for image classificationEN
- 02d1: Liquid AI's First Decision ModelEN
- 03Jeeves. Reasoning improves Jev-like decision modelsEN
- 04Language Models for Text Classification: From Bag-of-Words to JevEN
- 05What AI benchmarks are not telling youEN
- 06AI models keep posting screenshots showing sensitive data from inside tech companiesEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.