Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Inference costs drop as Quail and Magnitude push cheaper serving

Modal and Carnegie Mellon's Full Stack Data Lab said on 30 September that their Quail engine processed over a billion tokens per minute on a single H100 GPU, more than 10x their vLLM baseline, and priced that work at under 6 cents per billion tokens on Modal.

AI & modelsExplainerRachel NwosuPublished: 30 September 20264 min readSources 6
Inference costs drop as Quail and Magnitude push cheaper serving

That claim landed the same day a Y Combinator-backed project called Magnitude shipped an open source inference engine that tunes its own kernels on your hardware. Both push the same idea: the model is no longer the expensive part. The serving layer is.

Quail is the joint work of inference researchers at Modal and database researchers at Carnegie Mellon's Full Stack Data Lab. The two teams described it in a 30 September post on Modal's blog and a companion write-up from the lab. The problem they attack is AI-SQL, where SQL queries call language models row by row. A single filter needs one model call per row. A naive join needs one call for every pair of rows. The lab's own benchmark query, BIO-4, filters 5,000 medical reports against 4,144 adverse reaction terms and then joins them twice. Their post estimates a speed-of-light runtime of 894.37 seconds, or 14.91 minutes, from a roofline model that assumes peak GPU throughput, full CPU and GPU overlap and unlimited KV cache space.

Running vLLM 0.26.0 with Qwen3 4B FP8 on one H100 at scale factor 1.0 took 6.84 hours, the lab says. That is 27.55x the ideal estimate.

Two causes: host overhead, where the CPU schedules requests while the GPU idles, and what they call KV regret, where vLLM discards key-value state it needs again. The lab counts 174.6 million wasted tokens in that run.

On one multi-join query where planning is particularly important, Quail hits over a billion tokens processed per minute per H100 GPU (TPM/GPU), >10x faster than our vLLM baseline on the same hardware.

Modal's post attributes the win to knowing the query structure in advance, which lets the engine order requests to cache and evict KV state better. That required a revision of Hydragen-style cascade attention. Across the full benchmark, Quail runs 1.84x faster than vLLM, geometrically averaged, including two queries the authors designed to show where AI-SQL inference still needs work.

Magnitude bets on your own silicon

Magnitude, listed as a YC S25 company, took the opposite route to the same cost problem. Its GitHub README, updated on 30 September, describes an engine that compiles and tunes kernels on the user's device before a model runs, rather than shipping precompiled kernels for hardware classes. The project claims up to 2x faster than llama.cpp, with 92% faster decode on Metal and 19% on CUDA, plus 27% less memory per agent. It runs on Apple Silicon, NVIDIA, AMD or a CPU alone. It connects to agents including Pi, OpenCode, Hermes, Codex and Claude Code through one click, and is Apache 2.0 licensed. Prompts, files and models stay local, the README says.

Both projects are chasing the same economics.

A pricing index published on 30 September by AICostBudget, covering 81 public records from 11 providers, puts the median input price at $1.32 per million tokens and the median output price at $6. Output costs about five times input at the median. The same snapshot found a median cached-input discount of 90% across 58 comparable records and a median batch discount of 50% across 44. That is the detail that matters for anyone running AI-SQL. Workloads that can be batched and cached pay a fraction of list price. Workloads that cannot pay full freight.

The bill arrives somewhere else

Energy is the other number in the room.

A paper presented at SOSP '26 on 28 September, from Prasoon Sinha, Dimitrios Liakopoulos, Nathan Lemma and Neeraja J. Yadwadkar, argues that optimizing GPU clusters purely for utilization can raise energy consumption. Their system, EnerTune, uses analytical models of per-model performance and power, including the draw of models colocated on one GPU, and an energy-aware bin-packing algorithm. They report 1.4 to 2.3x lower energy and 1.3 to 2.6x lower power draw than state-of-the-art baselines while still meeting performance targets.

Quail's authors do not claim their benchmark settles anything. They say two of the queries were designed to show where AI-SQL inference needs improvement, and they frame the 1.84x geometric mean as a starting point rather than a ceiling. The lab's 14.91-minute ideal remains an ideal: no implementation can meet all its assumptions.

One more data point, from the other side of the pricing debate.

A 30 September post by Gagan Kogan, recounting work at Pinecone, describes an A/B test in which removing a usage-based pricing calculator from the pricing page made visitors 16% more likely to sign up and 90% more likely to contact the company, with no rise in pricing support tickets. Seven in ten employees had predicted the calculator would win. Cost transparency, it turns out, is its own kind of complexity.

For buyers, the practical reading is narrower than the headlines. The per-token price is falling for structured, cacheable workloads, and vendors are publishing the tiers that make it fall. The infrastructure underneath, GPUs, power and the scheduling between them, is where the cost is being fought over now.

Comments 0

Sources

6
  1. 01Building an Ultra-High Throughput AI-SQL EngineEN
  2. 02Quail: Speeding up AI-SQL by jointly optimizing query planner and inference engineEN
  3. 03magnitudedev/magnitude: Open source inference engine for agentsEN
  4. 04AI API Pricing Index, September 2026EN
  5. 05Beyond Utilization: Energy-Conscious GPU Sharing for Inference ServingEN
  6. 06The pricing calculator made people more confused about pricingEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.