Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Small decision models move on-device: Liquid AI's d1, Jev clones and a 2B vision classifier

Liquid AI published documentation for d1, its first "decision model", on 29 September, part of a fast-growing family of small on-device models that return calibrated probabilities instead of generated text.

AI & modelsNewsRachel NwosuPublished: 29 September 20265 min readSources 12
Small decision models move on-device: Liquid AI's d1, Jev clones and a 2B vision classifier

The docs describe three question types: Noul, a yes/no question returning a probability between 0 and 1; Choice, a pick-one-from-many question returning a distribution over named options; and Score, an ordered rubric returned as a probability-weighted position. The example call sends a customer complaint to model "d1:free" and returns 0.999 on the question "Is this message a complaint from the customer?"

That is a different shape of model to a chatbot. Nothing is generated. Liquid's documentation says a decision model "evaluates a situation and returns calibrated probabilities across a fixed set of outcomes in a single call with zero generated tokens", and that multiple questions can be evaluated against the same state in one request.

Why this class of model exists

The same pattern shows up across the newest sources in the dossier. Sebastian Raschka published a long technical explainer on 29 September tracing text classification from bag-of-words through RNNs, CNNs and transformers to what he calls "Jev-like APIs", and argues the appeal is cost and speed rather than raw accuracy. "Jev's advantage is that it can handle those classification tasks much faster and more cheaply," he writes, while warning that for a narrow, well-defined problem a purpose-built classifier will still win on all three counts.

PostHog's Jeeves repository, also published on 29 September, puts numbers on the trade-off. Jeeves is a 9B Jev-like model built on Qwen3.5-9B with LoRA and a pointer head, trained with SFT and CISPO to reason before it decides. The repo reports 0.889 on its own held-out test split against 0.857 for Jev and 0.822 for Kev-9B, and 0.935 on JevBench's public tiers against 0.866 for Jev. Inference runs at about 0.3 seconds per request without thinking and a 3.3 second median with it, on one H100 at fp8. It also runs on Apple Silicon in bf16 or FP8, and requires 48GB or more.

Evan Schwartz, who runs the personalised content feed Scour, published a request on 29 September for prompt caching on this class of API. His question set now has 54 questions (50 nouls, 3 choices, 1 score) at roughly 2,150 tokens, of which about 260 tokens are fixed overhead. Sent verbatim with every request across roughly 1.1 million documents a month, the questions are about 88% of his input tokens. His bill is under $150 a month, but he says caching would make batch workflows cheaper still.

The on-device end: a 2B vision classifier

The clearest small-model release in the batch is Qevi-2B, published to Hugging Face on 29 September. It is a full fine-tune of Qwen3-VL-2B-Instruct that answers closed questions about images by reading the model's own logits rather than generating text. Ask it whether there is a ladder in an image and it returns P(Yes) = 0.97, not a sentence.

The model card reports in-domain accuracy rising from 0.855 for the base model to 0.977, held-out domain accuracy from 0.745 to 0.889, and expected calibration error on held-out data falling from 0.160 to 0.054. Inference is listed as up to roughly 21 times faster when asking many questions about one image. The card is explicit that the model must be used through a logit readout and not .generate(), and that all three supported question types can share a single encoding of the image, isolated by a block-diagonal attention mask, with answers verified bitwise against asking them separately.

One evaluation in the dossier cuts against the enthusiasm. Casco, a security company, ran 2,449 findings through eight models including Jev, Claude and GPT and compared their CVSS 3.1 scores against its own published scores. Every model scored higher than the reference on average. Jev assigned 550 critical scores to a dataset containing 55 published critical findings, and 500 of those 550 were below critical in Casco's reference. On mean absolute error, GPT-6 Astra ranked first at 1.92 points and Jev seventh at 3.40. Casco's write-up is dated 28 September.

Elsewhere in the same 29 September batch, the small-model theme shows up in less obvious places. A free preview model listed on OpenRouter on 24 September as space-bunny-alpha claims a million-token context and reports using about one third of Qwen3.8 Flash's output tokens on the same benchmarks, 305,989 against 913,989, according to its own site. Simple AI launched Tango, an audio-native model for end-of-turn detection, claiming a median decision 300ms before a conventional voice-activity detector would spot a pause, and 15 missed turn endings out of 400 on the public LiveKit benchmark against 54 for Soniox and 197 for Deepgram Flux. Those are vendor-run comparisons on a public benchmark, not independent audits.

Two other releases in the window push the same direction: Jeb, a GitHub project that turns any OpenAI-compatible API into a decision model, and DesktopBrain, an on-device file organiser for Mac. Neither publishes benchmark numbers in the material available.

The pattern is consistent. The interesting small models of the last 72 hours are not miniature chatbots. They are narrow, calibrated, and cheap to run in places where a frontier model would be absurd: a phone, a laptop, a 2B vision encoder, a classifier that answers 54 questions for less than $150 a month. Whether they are accurate enough for the work they are pointed at is a separate question, and Casco's result suggests the answer is not always yes.

Comments 0

Sources

12
  1. 01d1: Liquid AI's First Decision ModelEN
  2. 02Qevi-2B: A Jev-style finetuned model for image classificationEN
  3. 03Language Models for Text Classification: From Bag-of-Words to JevEN
  4. 04Jeeves. Reasoning improves Jev-like decision modelsEN
  5. 05Every model (incl. Jev) we tested inflates security finding severityEN
  6. 06Please add prompt caching to Jev-style modelsEN
  7. 07Space-bunny-alpha a reasoning model from an anonymous providerEN
  8. 08Tango: Simple AI's Conversational Awareness ModelEN
  9. 09Jeb: Turn any OpenAI API into a decision modelEN
  10. 10DesktopBrain - AI File Organizer for Mac and All Folders - Private On-Device AIEN
  11. 11Language Models Act on Hidden ValenceEN
  12. 12Needed 1+1, built a functional programming languageEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.