Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Small decision models move on-device as Jev-style runners proliferate

Liquid AI published documentation for d1, its first decision model, on 29 September. The API returns calibrated probabilities across fixed outcomes in a single call, with zero generated tokens. The same day, a Hugging Face release showed a 2B vision model reading its own logits on-device.

AI & modelsAnalysisRachel NwosuPublished: 29 September 20264 min readSources 10
Small decision models move on-device as Jev-style runners proliferate

Liquid AI's d1 documentation went up on 29 September. Every query falls into one of three typed forms: a noul, a yes/no question that returns a probability between 0 and 1; a choice, a pick-one distribution over named options; and a score, a probability-weighted position on an ordered rubric. The docs give an example where a customer complaint returns 0.999 on a noul gate, and note that one request can carry several questions evaluated at once. That is the shape the on-device crowd has been waiting for.

On a phone or a laptop, the argument is arithmetic. A model that emits one number instead of a paragraph skips the decode loop, so latency stops scaling with answer length. It also makes the output thresholdable. A probability of 0.92 is a branch condition, not something a parser has to interpret.

Qevi-2B, released on Hugging Face on 29 September, is the clearest small-model example. It is a full fine-tune of Qwen3-VL-2B-Instruct that answers typed, closed questions about images by reading the LM-head logits at the reply position instead of generating text. The model card reports in-domain accuracy of 0.977 against 0.855 for the base Qwen3-VL-2B, held-out domain accuracy of 0.889 against 0.745, and expected calibration error on held-out data of 0.054 against 0.160.

The card is blunt about the catch. The numbers only reproduce through a logit readout, roughly 30 lines of ordinary transformers code, and calling .generate() will not match them, because that is not the path the model was fine-tuned on. The model is not a chat model either. Ask whether there is a ladder in the image and it returns P(Yes) = 0.97, not a sentence. The claimed payoff is up to roughly 21x faster inference when asking many questions about one image. Three separate questions can share a single encoding and be isolated with a block-diagonal attention mask, verified bitwise against asking them separately. Training covered 57,000 noul questions and 28,500 choice questions.

Reasoning puts the small models back in the race

The obvious objection is accuracy. Jev-style classifiers are cheap and calibrated, but as one implementation notes, they are weak at low accuracy, which is why pipelines keep a reasoning model as a fallback. PostHog's Jeeves repo, pushed on 29 September, attacks that directly. It is a 9B Jev-like model (Qwen3.5-9B, LoRA, pointer head) that thinks before it decides, trained with SFT and CISPO, plus a block-4 diffusion drafter. The published table puts Jeeves at 0.889 on held-out test data against 0.857 for Jev and 0.822 for Kev-9B, and 0.935 against 0.866 for Jev on JevBench's public tiers. It reports about 0.3 s per request without thinking and a 3.3 s median with it on one H100 at fp8, and it runs on Apple Silicon in bf16 or FP8.

Sebastian Raschka's 29 September history of text classification frames the whole wave. He traces the path from bag-of-words and naive Bayes through RNNs and transformers to what he calls Jev-like APIs, and argues Jev's advantage is speed and cost on classification rather than raw capability. Evan Schwartz, writing the same day, makes the deployment case concrete. He runs 54 questions over about 1.1 million documents a month on Scour, his questions are roughly 88% of input tokens, and his total bill stays under $150 a month. He asks vendors for prompt caching or reusable question sets.

Not everyone is convinced the outputs are trustworthy. Casco tested eight models on 2,449 security findings against its own published CVSS 3.1 scores and found every model overestimated severity. Jev assigned 550 critical scores to a dataset containing 55 published critical findings, with a mean signed error of +3.30 points. GPT-6 Astra was the most accurate at 1.92 points mean absolute error, with Jev seventh at 3.40. The same overdiagnosis risk applies to any on-device gate that fires on a threshold.

Infrastructure is following. Microsoft's developer blog argued on 29 September that public coding benchmarks do not predict performance on your own codebase. That is the same argument for evaluating a small decision model on your own inputs. JuliaHub reported the same day that swapping the agent harness, holding the model fixed, moved scores from 0.533 to 0.899, more than doubling the gap it measured between four frontier models. Tuneloop put a number on the evaluation itself: 30 replayed tasks and about $55 pin the gap to plus or minus 6 points, which is enough to justify routing routine work to a cheaper model at 2.5% of the cost. For an on-device classifier, that is the test that decides whether the small model ships.

Comments 0

Sources

10
  1. 01d1: Liquid AI's First Decision ModelEN
  2. 02Qevi-2B: A Jev-style finetuned model for image classificationEN
  3. 03Jeeves. Reasoning improves Jev-like decision modelsEN
  4. 04Language Models for Text Classification: From Bag-of-Words to JevEN
  5. 05Please add prompt caching to Jev-style modelsEN
  6. 06Every model (incl. Jev) we tested inflates security finding severityEN
  7. 07What AI benchmarks are not telling youEN
  8. 08The Best AI Models Fail at Physics: Coding Harnesses are to BlameEN
  9. 09How many tasks does it take to trust a cheaper model?EN
  10. 10Tango: Simple AI's Conversational Awareness ModelEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.