Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Decision models ditch text generation as small on-device AI shifts to logit readouts

Two model releases on 29 September pushed the same idea: instead of generating tokens, small models should return calibrated probabilities. Liquid AI's d1 and a Qwen3-VL-2B fine-tune called Qevi-2B both ship that way.

AI & modelsAnalysisRachel NwosuPublished: 29 September 20264 min readSources 6
Decision models ditch text generation as small on-device AI shifts to logit readouts

Both landed within six minutes of each other on Tuesday. Liquid AI published docs for d1 at 20:09 UTC, calling it its first decision model. MeerDevelopment pushed Qevi-2B to Hugging Face at 20:16 UTC. Neither one generates text.

Liquid AI's documentation is blunt about the mechanism. A decision model, it says, "evaluates a situation and returns calibrated probabilities across a fixed set of outcomes in a single call with zero generated tokens." The API exposes three question shapes: noul (yes/no, returns a float), choice (pick one from named options) and score (a probability-weighted position on an ordered rubric). In the docs' example, the string "I have been waiting over three weeks for my order and nobody has responded to my emails" returns a noul of 0.999 for the question "Is this message a complaint from the customer?" The response also reports usage.output_tokens as 0.

Read the logits, skip the sentence

Qevi-2B takes a different route to the same destination. It is a full fine-tune of Qwen3-VL-2B-Instruct that answers closed questions about images by reading the LM-head logits at the position where the model would start replying, restricted to allowed answer tokens. Ask it whether there is a ladder in a photo and it returns P(Yes) = 0.97. The model card warns that calling .generate() and parsing the text will not reproduce the published numbers, "because that is not the path the model was fine-tuned on."

The claimed gains are large and specific. Against the base Qwen3-VL-2B run through the identical readout path, Qevi-2B scores 0.977 in-domain accuracy versus 0.855, and 0.889 on held-out domains versus 0.745. Expected calibration error on held-out data drops from 0.160 to 0.054. The card also reports roughly 21x faster inference when asking many questions about one image, because question packing shares a single image encoding.

One caveat deserves a close read. The training corpus covers only noul and choice questions, so the score type is untested. And the card explicitly retracts earlier temperature recommendations: at T = 2.45 the held-out ECE returns to 0.160, wiping out the entire calibration gain.

Small, fast, and pointed at classification

This is the on-device argument in its current form. Ailoitte's guide to small model development, published in the last 12 hours, frames it in cost terms: serving a 7B model can be "10 to 30 times cheaper in latency, energy, and compute" than a 70B to 175B model, citing a 2025 NVIDIA position paper. It also notes that small models run on-premises, which matters for healthcare and finance data.

Sebastian Raschka's 29 September essay on text classification makes the same point from the other direction. He describes Jev, a model that caused "quite a cultural phenomenon in technical communities in the past 2 weeks," as essentially a text classifier: faster and cheaper than general LLMs at classification, yet more general than task-specific models. He is careful to note he has no affiliation with Jev and no free access to it.

PostHog's jeeves repository, updated on 29 September, adds reasoning to that pattern. Jeeves is a 9B Jev-like model built on Qwen3.5-9B with LoRA and a pointer head, trained with SFT and CISPO. It reports 0.889 overall on a held-out test split against Jev's 0.857 and Kev-9B's 0.822, and 0.935 on JevBench's public tiers against Jev's 0.866. Thinking costs time: about 3.3 seconds median per request on one H100 in fp8, against 0.3 seconds without it.

The numbers do not point one way. Jeeves beats Jev overall but loses on Transfer (0.746 versus 0.800) and on MMLU-Pro 10-way (0.739 versus 0.840). It also answers 5.5% of unknowable questions at p ≥ 0.9, better than Jev's 9.0% but worse than Kev-9B's 0.0%.

Market forecasts for the category vary widely, which is worth flagging rather than averaging. MarketsandMarkets puts the SLM market at USD 0.93 billion in 2025 rising to USD 5.45 billion by 2032 at a 28.7% CAGR, while wizr.ai cites a figure of USD 7.76 billion in 2023 growing at 15.6% from 2024 to 2030. The two estimates describe different-sized markets.

What the pair of 29 September releases suggest is narrower and more concrete: for classification-shaped work, the output format itself is being redesigned around probabilities, not prose.

Comments 0

Sources

6
  1. 01d1: Liquid AI's First Decision ModelEN
  2. 02Qevi-2B: A Jev-style finetuned model for image classificationEN
  3. 03Jeeves. Reasoning improves Jev-like decision modelsEN
  4. 04Language Models for Text Classification: From Bag-of-Words to JevEN
  5. 05Small Language Model Development: Lower Cost, More PrivacyEN
  6. 06The Rise of Small Language Models in Enterprise AI Adoption [2026]EN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.