Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Small decision models move on-device as Liquid AI, PostHog and Qevi ship classifiers

Liquid AI published documentation for d1, its first decision model, on 29 September. The same day, PostHog released Jeeves, a 9B Jev-style classifier, and a Hugging Face card appeared for Qevi-2B, a 2B image classifier.

AI & modelsNewsRachel NwosuPublished: 29 September 20268 min readSources 6
Small decision models move on-device as Liquid AI, PostHog and Qevi ship classifiers

Three small, task-specific models landed within hours of each other on 29 September. All three work on the same idea: skip text generation, read probabilities straight out of a model, and run the thing on hardware you control.

Liquid AI's documentation page calls d1 "a new class of AI model purpose-built for structured decisions." A decision model does not generate tokens one at a time. Instead, the page says, it "evaluates a situation and returns calibrated probabilities across a fixed set of outcomes in a single call with zero generated tokens." The company's own example response shows usage.output_tokens: 0 for a complaint-detection query that returned a noul value of 0.999.

Three question shapes, no generated text

Liquid splits decision questions into three types. A noul is a yes or no question that returns a probability between 0 and 1, which the docs describe as "a boolean on a sliding scale." A choice returns a distribution over named options, such as {"billing": 0.65, "technical": 0.30, "account": 0.05}. A score returns a probability-weighted position on an ordered rubric, so a three-level urgency question could come back as 1.85, between "Medium" and "High".

The documentation is explicit that noul and score are not interchangeable. A noul at 0.5 means maximum uncertainty between yes and no, and "says nothing about degree or intensity." If the code needs to compare against a threshold on a continuum, the docs point to score. If it gates on a boolean, noul. The API takes a state plus one or more named questions, and answers all of them in a single call. Client libraries exist for Python (typesafe-sdk) and JavaScript (@typesafe-ai/sdk), and keys are prefixed with liquid_.

Liquid's page is documentation, not a benchmark paper, so it carries no accuracy figures for d1. That matters, because the same week produced two models that do publish numbers, and the gap between them is instructive.

Jeeves adds reasoning to the Jev formula

PostHog's GitHub repository for Jeeves describes it as "a reasoning Jev-style classifier with a diffusion drafter, trained with SFT and CISPO." It is a 9B model built on Qwen3.5-9B with LoRA and a pointer head, plus a block-4 diffusion drafter and the full training code and data. The repository claims it beats Kev-9B and Jev on test data it was never trained on, at 0.889 against 0.822 and 0.857, and on JevBench's public tiers at 0.935 against 0.866 for Jev.

The repository publishes a fuller table, and it is not uniformly flattering. Jeeves scores 0.746 on transfer tasks covering MMLU-Pro and buried state, below Jev's 0.800. On MMLU it drops to 0.793 against Jev's 0.900, and on MMLU-Pro 10-way it falls to 0.739 against 0.840. It wins on PAWS (0.875 against 0.788), held-out rule structures (1.000 against 0.885) and contrastive policies (1.000 against 0.963). It also answers unknowable questions at p of 0.9 or above less often than Jev, 0.055 against 0.090, where lower is better.

The repository also notes that without thinking, the same checkpoint scores 0.804 on its test split of 2,962 items, against 0.840 with thinking. Latency is the trade: about 0.3 seconds per request without thinking, and a 3.3 second median with it, on one H100 at fp8 precision. On an M4 Pro, one question thinks at roughly 20 tokens per second. The weights use 21 GB in bf16, and the default caches add another 28 GB, which is why the README recommends smaller cache settings on a 48 GB Mac.

One caveat is printed in the table itself. No Kev-9B JevBench result is published, and the two asterisked rows are Kev-8B on Qwen3, so part of that comparison is being made against a different model than the column header suggests.

Qevi-2B takes the same trick to images

The Qevi-2B model card on Hugging Face describes a full fine-tune of Qwen3-VL-2B-Instruct that answers typed, closed questions about images "by reading the model's own logits instead of generating text." Ask whether there is a ladder in a picture and it returns P(Yes) = 0.97, not a sentence. The card is blunt that it is not a chat model, and warns that calling .generate() and parsing the text will not reproduce the published numbers, because that is not the path the model was fine-tuned on.

The readout is about 30 lines of ordinary transformers code with no trust_remote_code. Against base Qwen3-VL-2B-Instruct run through the identical readout path, Qevi-2B goes from 0.855 to 0.977 in-domain accuracy and from 0.745 to 0.889 on 12 held-out domains. Expected calibration error on held-out data drops from 0.160 to 0.054. The card says the speed advantage reaches about 21 times when asking many questions about one image, because a block-diagonal attention mask lets several questions share one encoding of the image, verified bitwise against asking them separately.

The card warns against the temperature scaling earlier versions recommended: at T = 2.45 the held-out ECE is 0.160, essentially the base model's figure, and "the entire calibration gain is cancelled out."

There is a second honesty note. The training corpus contains only noul and choice questions. The engine supports score and the base model handles such questions zero-shot, but the fine-tune never saw one, so the card states that score accuracy and calibration are unmeasured and should be treated as untested.

Why the on-device angle keeps recurring

Sebastian Raschka's 29 September write-up on text classification history is useful here, because it frames what these models are actually giving up. He notes that the latest GPT and open-weight models can do the same classification tasks as Jev while being far more general, but that a Jev-style model's advantage is speed and cost. At the other end, "for a narrow, well-defined problem, Jev probably won't classify anything better, faster, or cheaper than a special-purpose classifier." The selling point is the middle ground: more general than a task-specific model, cheaper than a frontier one.

Raschka's article traces the lineage back to bag-of-words representations feeding naive Bayes, logistic regression, SVMs and XGBoost, and notes the Gmail spam filter was allegedly a naive Bayes model over bag-of-words. The difference now is that the classifier is a fine-tuned language model with a calibrated probability head, and it can be shipped as weights rather than called as a service.

That is where the on-device argument bites. Jeeves runs on CUDA in bf16 or fp8, and also on Apple Silicon via MPS, with an fp8 build that shrinks weights to 11.5 GB and raises thinking throughput on an M4 Pro from about 20 to about 31 tokens per second. The repository says accuracy and NLL did not change measurably on dev questions, and a pre-quantized PostHog/jeeves-fp8 repo is published so the download is roughly half the size.

Qevi-2B is smaller still, at 2B parameters, which is the kind of footprint that fits a laptop or a phone rather than a rack. Liquid's docs describe the model as evaluating state and questions in one call with zero output tokens, which is the property that makes per-request cost predictable.

The measurements are still thin

None of this is free of problems. The most recent adjacent research, an arXiv paper submitted on 28 September by Cameron Berg and Caspar Kaiser, found that across seven open-weight models from five families, hidden valence-related activation patterns predictably govern later choices, even when every visible token is identical. The authors write that whether these traces are accompanied by any subjective experience relevant to model welfare "remains unclear." It is a reminder that reading a model's internal state is now a design pattern, not just a diagnostic.

The practical objection comes from Microsoft's developer blog, published on 29 September, which argues that public benchmarks tell you little about your own workload. The post cites Goodhart's law and points out that SWE-bench tasks come from public repositories, so overlap with training data is inevitable and grows each generation. A model that scores 92% on SWE-bench, the post says, is demonstrably good at resolving well-documented issues in popular repositories, and that says nothing about your internal library.

The same logic applies to decision models. Jeeves' own table shows it losing to Jev on MMLU and MMLU-Pro while winning on held-out rule structures, which means the answer to "which is better" depends entirely on which distribution your tickets, images or documents come from. Liquid publishes no accuracy figures for d1 at all. Qevi-2B's card lists score as untested. The numbers that do exist are mostly self-reported by the teams shipping the models, on splits they chose.

What is new this week is not that small classifiers exist. It is that three separate teams shipped them with the same interface idea, calibrated probabilities over a fixed answer set in one forward pass, and shipped them small enough to run without a datacentre. Whether that is enough to displace the frontier-model-with-a-fallback pattern the Jeeves README describes is a question the published benchmarks cannot answer yet.

Comments 0

Sources

6
  1. 01d1: Liquid AI's First Decision ModelEN
  2. 02Jeeves. Reasoning improves Jev-like decision modelsEN
  3. 03Qevi-2B: A Jev-style finetuned model for image classificationEN
  4. 04Language Models for Text Classification: From Bag-of-Words to JevEN
  5. 05Language Models Act on Hidden ValenceEN
  6. 06What AI benchmarks are not telling youEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.