Inference Engine Quail Hits 1B Tokens a Minute on One H100
Two research teams published new inference-engine results on 30 September, claiming AI-SQL queries can run more than ten times faster than a vLLM baseline on the same hardware.

On 30 September, Modal and Carnegie Mellon University's Full Stack Data Lab published results for Quail, a query-aware inference layer. The team says it processes over a billion tokens a minute per H100 GPU on a single multi-join query. That is more than 10x faster than their vLLM baseline on the same hardware. The claim comes with a cost figure attached: under 6 cents per billion tokens on Modal.
The work targets AI-SQL, the extension of SQL that lets a query call a language model row by row. The Full Stack Data Lab blog describes a filter that needs one LLM call per row, and a naive join that needs one call for every pair of rows in its two input tables. That arithmetic gets large fast. In a benchmark query the lab calls BIO-4, 5,000 medical reports are joined against 4,144 reaction terms, with the reaction list used twice. The lab writes that a vLLM 0.26.0 baseline with Qwen3 4B FP8 on one H100 took 6.84 hours at scale factor 1.0. It calculates that as 27.55x an optimistic roofline estimate of 14.91 minutes.
The lab attributes the gap to two things.
The first is host overhead: the CPU spends long stretches scheduling requests while the GPU waits. The second it calls KV regret. vLLM discards key and value state it needs again later, so the BIO-4 run processes 174.6 million tokens it did not have to.
"The big win is that with a structured query in hand, you can order requests to better cache (and evict) KV," the Modal blog states.
Quail is not alone in chasing this. Magnitude, a Y Combinator S25 company, opened its repository on GitHub on 30 September with a different pitch: an open source inference engine for agents that compiles and tunes kernels on the user's own device. Its README claims open models run up to 2x faster than llama.cpp, with 92% faster decode on Metal and 19% on CUDA, and 27% less memory per agent. It lists Apple Silicon, NVIDIA, AMD and plain CPUs as targets, and ships under Apache 2.0. Both projects argue that the general-purpose serving stack leaves performance on the table. Quail's answer is knowing the query structure in advance. Magnitude's answer is knowing the silicon.
The economics behind that argument are visible in the pricing data. AICostBudget's September index, frozen at a 30 September UTC cutoff, covers 81 public pricing records across 11 providers. Across 69 records with a current input-token scalar, prices range from $0.10 to $30 per 1M tokens. Output runs from $0.10 to $180. The median output-to-input ratio across the 69 records carrying both scalars is 5.0x. Discounts matter as much as list prices. Over 58 comparable records, the index reports a median cached-input discount of 90%. Over 44 batch-capable records, the median batch input discount is 50%. The methodology page is explicit that regional pricing and negotiated rates are not captured, and that these are planning inputs rather than invoices.
Hardware-side research is moving in the same direction. A paper published at SOSP '26 on 28 September presents EnerTune, a serving system that uses analytical models of per-model performance and power, plus the power draw of colocated models on shared GPUs, to drive an energy-aware bin-packing algorithm. The authors, Prasoon Sinha, Dimitrios Liakopoulos, Nathan Lemma and Neeraja J. Yadwadkar, report energy reductions of 1.4-2.3x and power-draw reductions of 1.3-2.6x over state-of-the-art baselines while meeting performance SLOs.
Their starting point is that GPU clusters remain heavily underutilized, and that optimizing purely for utilization can raise energy consumption. At scale, they argue, profiling every model across hundreds of configurations is prohibitively expensive, so estimation has to replace measurement.
Not every recent result is about the serving stack. MindOn introduced Mind-1 on 30 September, a physical AI model it says cuts robot inference latency from 82ms to 32ms, bringing manipulation tasks to human-level cycle times. The company's post notes that at an end-effector speed of 3 m/s, 10 ms of latency equals roughly 3 cm of movement, enough to affect precise grasping and insertion. That is the same latency budget argument, applied to a robot arm rather than a data warehouse.
One result cuts the other way. A post by Gleb Kogan published on 30 September describes Pinecone removing a usage-based pricing calculator from its pricing page. In an A/B test, visitors who did not see the calculator were 16% more likely to sign up and 90% more likely to contact the company, with no rise in pricing support tickets. The internal poll before the test found 7 of every 10 employees expected the calculator version to win. Cost tooling, in other words, can mislead as easily as it informs, which is a caution for anyone reading per-token benchmarks as a bill.
The Quail numbers remain vendor-published research. The 1.84x figure the team gives for its newly released benchmark, geometrically averaged over tasks, is a much more modest claim than the billion-tokens-per-minute headline. The benchmark itself includes two queries the authors say they designed to show where AI-SQL inference still needs work.
Sources
7- 01Hitting 1B tokens/minute on 1 GPU combining a query planner and inference engineEN
- 02Building an Ultra-High Throughput AI-SQL EngineEN
- 03Launch HN: Magnitude (YC S25) - Self-optimizing inference engine for agentsEN
- 04AI API Pricing Index - September 2026EN
- 05Beyond Utilization: Energy-Conscious GPU Sharing for Inference ServingEN
- 06Mind-1: Physical AI at Human Speed, Built for Real WorkEN
- 07The pricing calculator made people more confused about pricingEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.