Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Small models move on-device: Liquid AI ships d1, Qevi-2B reads logits instead of talking

Liquid AI published documentation for d1 on 29 September. The docs describe a decision model that returns calibrated probabilities across a fixed set of outcomes in a single call, with zero generated tokens. The same day, a 2B vision model called Qevi-2B appeared on Hugging Face with a logit readout instead of text generation.

AI & modelsAnalysisRachel NwosuPublished: 29 September 20264 min readSources 12
Small models move on-device: Liquid AI ships d1, Qevi-2B reads logits instead of talking

Liquid AI's d1 docs landed on 29 September. They describe a class of model that never writes a sentence. It takes one state (text or a JSON object) plus one or more typed questions, and returns probabilities. Three question types are supported: noul for yes/no, choice for pick-one-from-many, and score for an ordered rubric. A three-level question can come back as 1.85, between Medium and High. "Decision models are a new class of AI model purpose-built for structured decisions," the page says.

That framing is not unique to Liquid. Tuesday 29 September also brought Qevi-2B, a full fine-tune of Qwen3-VL-2B-Instruct published by MeerDevelopment on Hugging Face. It answers closed questions about images by reading the model's own logits. Ask it whether there is a ladder in a photo and it returns P(Yes) = 0.97, not a sentence. The model card reports in-domain accuracy of 0.977 against 0.855 for the base Qwen3-VL-2B, held-out domain accuracy of 0.889 against 0.745, and expected calibration error on held-out data of 0.054 against 0.160. It also claims up to roughly 21x faster inference when asking many questions about one image.

Why the small ones are the interesting ones

The economics matter more than the architecture. Evan Schwartz runs a personalised content feed called Scour. He wrote on 29 September that he sends 54 questions (50 nouls, 3 choices, 1 score) against about 1.1 million documents a month. His question set is roughly 2,150 tokens and represents about 88% of input tokens per request, because the questions are resent verbatim every time. His total bill is under $150 a month. He says a general-purpose LLM could not match that on that volume. He is asking the labs to add prompt caching or reusable question sets, and points to an open request on the TypeSafe SDK repository.

Other people are already building workarounds. A project called Jeb, posted to GitHub on 29 September, turns any OpenAI-compatible endpoint into a Jev-style decision service by reading token log probabilities and returning structured JSON at POST /v1/systemone. Its README is blunt about the dependency: an API can be OpenAI-compatible for ordinary chat while lacking the logprobs Jeb needs, so you have to check that capability before picking a provider. PostHog published Jeeves the same day, a 9B Qwen3.5-9B model with LoRA and a pointer head that reasons before it decides, plus a block-4 diffusion drafter. It reports 0.889 on its own held-out test split against 0.857 for Jev and 0.822 for Kev-9B, and 0.935 on JevBench public tiers against 0.866 for Jev. The 0.3 s per request figure is without thinking. With thinking the median is 3.3 s on one H100 at fp8.

One cautionary data point arrived from Casco on 28 September. The security firm ran 2,449 findings through eight models, including Jev, Claude and GPT, asking each to pick the eight CVSS 3.1 base metrics. Every model scored higher than Casco's published reference on average. Jev assigned 550 critical ratings to a dataset containing 55 published critical findings. GPT-6 Astra had the lowest mean absolute error at 1.92 points; Jev ranked seventh at 3.40. Structured output did not translate into better severity judgment.

Elsewhere on 29 September, Sebastian Raschka published a long technical history of text classification that tries to place Jev in context. He notes that bag-of-words with naive Bayes goes back to at least 1961 and that even Gmail's original spam filter allegedly used it. His assessment is measured: for a narrow, well-defined problem, a Jev-style model probably will not beat a special-purpose classifier, but it is far more general. Simple AI also used the day to launch Tango, an audio-native model for end-of-turn detection. The company says it lands a median 300 ms before a conventional voice-activity detector, and fires 3.7 to 4.7 times less often mid-speech than Smart Turn at equal turn-end coverage. On the public LiveKit benchmark it misses 15 of 400 true turn endings, against 54 for Soniox and 197 for Deepgram Flux.

The through-line is that the on-device and small-model tier is where deployment is actually happening, and where the measurement arguments are now being had. Qevi-2B's card is explicit that its numbers only reproduce through the logit readout, not through .generate(): "If you call .generate() and parse the text, you will not reproduce them, because that is not the path the model was fine-tuned on." That is a useful warning for anyone benchmarking these systems. A model can look worse than it is simply because it is being asked to talk.

Comments 0

Sources

12
  1. 01d1: Liquid AI's First Decision ModelEN
  2. 02Qevi-2B: A Jev-style finetuned model for image classificationEN
  3. 03Please add prompt caching to Jev-style modelsEN
  4. 04Jeb: Turn any OpenAI API into a decision modelEN
  5. 05Jeeves. Reasoning improves Jev-like decision modelsEN
  6. 06Every model (incl. Jev) we tested inflates security finding severityEN
  7. 07Language Models for Text Classification: From Bag-of-Words to JevEN
  8. 08Tango: Simple AI's Conversational Awareness ModelEN
  9. 09The System One models ecosystemEN
  10. 10AI models keep posting screenshots showing sensitive data from inside tech companiesEN
  11. 11OpenAI announces 'dots' agent after scrapping launch of new AI model over safety concernsEN
  12. 12Language Models Act on Hidden ValenceEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.