Inference Cost Race: Quail Claims 1B Tokens/Minute, Magnitude Tunes Kernels on Device
Two engineering teams published separate claims on 30 September that the cost of running open models can be cut by rewriting how inference engines handle hardware and query structure, not by waiting for cheaper GPUs.

Modal and Carnegie Mellon University's Full Stack Data Lab said on 30 September that Quail, the inference engine they built together, processed more than a billion tokens per minute on a single H100 GPU for one multi-join AI-SQL query. That is more than 10x their vLLM baseline on the same hardware, according to the Modal blog post. On Modal's own cloud it works out to under 6 cents per billion tokens.
The claim rests on one specific workload. AI-SQL extends ordinary SQL with functions that call a language model, so a filter can trigger one model call per row and a join one call per row pair. The Full Stack Data Lab's technical write-up describes a benchmark query, BIO-4, over 5,000 medical reports and 4,144 reaction terms. Running it on vLLM 0.26.0 with Qwen3 4B FP8 on one H100 took 6.84 hours at scale factor 1.0. The authors calculate that this is 27.55x a roofline estimate of 14.91 minutes.
Engine work, not silicon
That gap is the argument. The authors blame host overhead, with the CPU scheduling requests while the H100 sits idle, and what they call KV regret: vLLM discarding key and value cache it needed again later. The lab's blog says the query processed 174.6 million tokens of wasted KV work. Quail takes the query plan as an input instead, which lets it order requests to reuse cache and evict it deliberately.
The two write-ups do not stress the same numbers. Modal frames the result around the single best query and the 1B tokens per minute figure. The Full Stack Data Lab reports that Quail runs 1.84x faster than vLLM when geometrically averaged across its new AI-SQL benchmark, including two queries the authors designed to expose remaining weaknesses. Both sets of numbers come from the same collaboration, so they are not independent verification.
On the same day, Magnitude, a Y Combinator S25 company, opened the source of its own inference engine under Apache 2.0. Its GitHub repository claims open models run up to 2x faster than llama.cpp, with 92% faster decode on Metal and 19% on CUDA. The mechanism differs from Quail's: Magnitude compiles and tunes kernels on the user's own device before a model runs, targeting Apple Silicon, Nvidia, AMD or a plain CPU. The repository also claims 27% less memory per agent and prefix cache sharing between concurrent sessions.
Both releases land in a market where list prices are not falling uniformly. AICostBudget's September snapshot, frozen at the 30 September UTC cutoff, covers 81 public pricing records across 11 providers. It puts median input at $1.32 per million tokens and median output at $6, a ratio of 5.0x, with cached input discounting at a median of 90% and batch pricing at 50%. Output, not prompt ingestion, is usually the more expensive component.
Hardware-level work is also being measured against energy rather than utilization. A paper published at SOSP '26 on 28 September, Beyond Utilization, argues that optimizing GPU multiplexing for utilization alone can raise energy consumption. Its system, EnerTune, uses analytical models of per-model power and the power draw of colocated models. The authors report energy reductions of 1.4x to 2.3x and power reductions of 1.3x to 2.6x against unnamed state-of-the-art baselines while meeting performance targets.
One cautionary note comes from outside infrastructure entirely. Pinecone's Gaurav Kogan described on 30 September how the company removed its usage-based pricing calculator after an A/B test. Visitors who did not see it were 16% more likely to sign up and 90% more likely to make contact, with no rise in pricing support tickets, he wrote. The lesson transfers: a tool meant to clarify cost can distort the decision it was built for.
Sources
6- 01Hitting 1B tokens/minute on 1 GPU combining a query planner and inference engineEN
- 02Building an Ultra-High Throughput AI-SQL EngineEN
- 03Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agentsEN
- 04AI API pricing index - September 2026EN
- 05Beyond Utilization: Energy-Conscious GPU Sharing for Inference ServingEN
- 06The pricing calculator made people more confused about pricingEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.