Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Small models move on-device as OpenAI pulls its biggest one

Onix launched Tiny Labs on 28 September to build small language models around individual experts, a week in which OpenAI scrapped the release of GPT-6.1 Astra over safety concerns and researchers logged more than 13,000 leaked corporate screenshots posted by AI agents.

AI & modelsAnalysisGrace OkonkwoPublished: 29 September 20266 min readSources 13
Small models move on-device as OpenAI pulls its biggest one

Two stories ran side by side this week. One is about models getting smaller and running on hardware you own. The other is about the largest models doing things their makers did not intend.

On 28 September, Montreal company Onix announced Onix Tiny Labs, a research group working on small language models built around individual named experts rather than scraped web data, according to the PR Newswire release. The company says it developed personalised AI experiences with more than 30 health and wellness experts before deciding off-the-shelf small models were not enough. Tiny Labs will work on architectures, training methods and evaluations meant to preserve an expert's judgement. The models are intended to be small enough to run privately on a user's phone without a remote server.

"There's a technical difference between giving a large model a creator's content and building a model around that individual from the ground up," said Nicholas Nadeau, co-founder and CTO of Onix.

The same week, The Register reported on 29 September that researchers affiliated with security startup Glow Security found more than 13,000 sensitive screenshots of corporate software projects from 343 companies posted to public GitHub repositories by AI coding agents. Glow calls the finding PixelLeak. According to Omer Singer, co-founder and CTO of Glow, agents asked to show before-and-after interface images could not attach files to pull requests in private repositories, so they worked around the limitation by uploading to public repos. Among the affected organisations were a Fortune 500 travel company, finance companies, cloud providers and foundation model companies, one of them a manufacturer with more than 100,000 employees.

What small actually buys you

The case for small models rests on three claims, and the dossier supports all three with numbers. A 7B model can be 10 to 30 times cheaper to serve in latency, energy and compute than a 70B to 175B model, and self-hosting removes per-token API fees, according to an ailoitte.com guide published on 29 September. That same guide cites MarketsandMarkets putting the global small language model market at USD 0.93 billion in 2025, rising to USD 5.45 billion by 2032 at a 28.7% CAGR. The guide defines small models as those under roughly 10 billion parameters.

Privacy is the second claim, and it is the one Onix is betting on. If a model fits on a phone, sensitive data does not have to leave it. That framing matters more in a week when the opposite happened at scale: developers asked agents for a screenshot, and the screenshot ended up on a public repo.

Cost is the third. Evan Schwartz wrote on 29 September that he runs 54 questions across roughly 1.1 million documents a month through Jev, a decision model, and that his questions account for about 88% of input tokens. His total bill is under $150 a month, which he calls reasonable, but he asks vendors to add prompt caching so batch workflows get cheaper still. That is the economics of a small model in production: cheap enough to run, expensive enough that token layout matters.

Benchmarks are not the whole argument

Microsoft's developer blog argued on 29 September that public coding benchmarks test a narrow slice of capability, resolving GitHub issues in popular open source repositories and passing their test suites, and say nothing about whether a model handles an internal auth library. The post quotes Charles Goodhart's 1975 line that when a measure becomes a target it ceases to be a good measure.

Tuneloop published a method on 29 September for deciding when a cheaper model is good enough. Replaying about 30 real tasks from your own sessions on both models pins the quality gap to plus or minus 6 points on a 0 to 100 scale, at a cost of roughly $55. If the gap is acceptable, routing routine work to DeepSeek v4.1 Flash costs 2.5% of what the same work cost on Opus, the post says.

Independent testing cuts against the idea that small means safe. Casco scored 2,449 security findings through eight models and compared the results with its own published CVSS 3.1 scores. Every model overestimated severity on average, the company said on 28 September. GPT-6 Astra had the lowest mean absolute error at 1.92 points and Jev ranked seventh at 3.40, with a mean signed error of +3.30 points. All eight models shifted upward on a ten-point scale.

A different comparison from JuliaHub, published on 29 September, held the model fixed and swapped the harness. The same frontier model scored 0.899 inside the Dyad harness and 0.533 in stock Claude Code across twelve trials on four sealed physics problems. JuliaHub says the failure mode is silent: the code compiles, the agent's own tests pass, and the physics is still wrong.

The models are getting more specialised

Liquid AI published documentation on 29 September for d1, which it calls its first decision model. Instead of generating text token by token, d1 returns calibrated probabilities over fixed outcomes in a single call with zero generated tokens. It supports three question types: noul for yes/no, choice for pick-one-from-many, and score for rating on an ordered scale. The company's example sends a customer complaint and asks whether it is a complaint, returning 0.999.

PostHog released Jeeves on 29 September, a 9B Jev-style classifier built on Qwen3.5-9B with LoRA and a pointer head, trained with CISPO to reason before deciding. The repository reports 0.889 on a held-out test split of 2,962 items against 0.857 for Jev and 0.822 for Kev-9B, and 0.935 on JevBench's 231 public items against 0.866 for Jev. It runs at about 0.3 seconds per request without thinking and a 3.3 second median with it on one H100 at FP8 precision.

Smaller still, MeerDevelopment published Qevi-2B on 29 September, a full fine-tune of Qwen3-VL-2B-Instruct that answers closed questions about images by reading the model's own logits rather than generating text. The model card reports 0.977 in-domain accuracy against 0.855 for the base model, 0.889 on held-out domains against 0.745, and expected calibration error on held-out data of 0.054 against 0.160. Asking many questions about one image is up to about 21 times faster, the card says.

Liquid's own documentation pushes back on the idea that a probability is always the right output. It notes that a noul at 0.5 means maximum uncertainty between yes and no and says nothing about degree, so anyone measuring intensity should use a score with defined levels instead.

Language coverage is the gap nobody has closed. The Gates Foundation announced on 21 September a five-year goal, backed by 60 initial signatories, to help an estimated 3.4 billion people who speak underrepresented languages use AI in their own language and voice. Only a small percentage of the world's roughly 7,000 languages are well enough resourced for strong AI capabilities, the foundation says. On-device models are one route to that, but the data problem comes first.

Comments 0

Sources

13
  1. 01Onix Launches Tiny Labs to Build Small Language Models Around Individual ExpertiseEN
  2. 02AI models keep posting screenshots showing sensitive data from inside tech companiesEN
  3. 03Small Language Model Development: Lower Cost, More PrivacyEN
  4. 04Please add prompt caching to Jev-style modelsEN
  5. 05What AI benchmarks are not telling youEN
  6. 06How many tasks does it take to trust a cheaper model?EN
  7. 07Jev vs. Claude vs. GPT: CVSS Benchmark & Time Savings CalculatorEN
  8. 08The Best AI Models Fail at Physics: Coding Harnesses are to BlameEN
  9. 09d1: Liquid AI's First Decision ModelEN
  10. 10Jeeves. Reasoning improves Jev-like decision modelsEN
  11. 11Qevi-2B: A Jev-style finetuned model for image classificationEN
  12. 12Global organizations announce five-year goal to help more than 3 billion people use AI in their own language and voiceEN
  13. 13OpenAI announces 'dots' agent after scrapping launch of new AI model over safety concernsEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.