Quail and Magnitude Attack Inference Costs From Opposite Ends of the Stack
Two open source projects published on 30 September claim the same prize from different directions. Quail reports over a billion tokens processed per minute on a single H100 by planning SQL queries against the inference engine. Magnitude says it runs open models up to 2x faster than llama.cpp by compiling kernels on the user's own hardware.

The pitch from Modal and Carnegie Mellon University's Full Stack Data Lab is specific. Quail, the QUery-Aware Inference Layer, hit over a billion tokens processed per minute per H100 GPU on one multi-join query, the team wrote on 30 September. That is more than 10x faster than their vLLM baseline on the same hardware. On Modal's cloud it works out to under 6 cents per billion tokens.
The number is large enough to look like a typo.
It is not a claim about a new model or a new chip. It is a claim about ordering.
What Quail actually does
AI-SQL extends SQL with functions that call a language model per row. A filter needs one model call per row. A naive join needs one call for every pair of rows across two tables. The Full Stack Data Lab blog, published 24 September and reposted by Modal, walks through a benchmark query called BIO-4: 5,000 medical reports joined against 4,144 reaction terms, twice. The lab's arithmetic puts the theoretical floor for that query at 894.37 seconds, or 14.91 minutes, assuming peak GPU throughput, full CPU-GPU overlap and unlimited space for retained KV cache. Running vLLM 0.26.0 with Qwen3 4B FP8 on one H100 took 6.84 hours, 27.55x that lower bound. The blog attributes the gap to two things. The first is host overhead: the CPU schedules requests while the H100 sits idle. The second is what it calls KV regret: vLLM discards key-value cache it needs again later. On BIO-4 at scale factor 1.0, the lab says vLLM processed 174.6 million, and the sentence trails off in the text we have.
Quail's fix is to use the query structure itself as a scheduling signal. If the engine knows the plan, it can order requests to reuse cache and evict it deliberately, rather than reactively. The Modal post describes this as a slight revision of Hydragen-style cascade attention, and credits the collaboration between inference researchers at Modal and database researchers at CMU.
Across a newly released AI-SQL benchmark, Quail runs 1.84x faster than vLLM geometrically averaged over tasks, a much smaller figure than the billion-tokens headline. The Modal post says the benchmark includes two queries designed to show where AI-SQL inference still needs work. That is a useful disclosure: the 10x number is a best case on one query, not a general speedup.
Magnitude goes the other way
Magnitude, a Y Combinator S25 company, posted its repository on 30 September with a different argument. Its inference engine compiles and tunes kernels on the actual device before a model runs, instead of shipping kernels precompiled for broad hardware classes. The README claims up to 2x faster than llama.cpp, with 92 percent faster decode on Metal and 19 percent on CUDA. It also claims 27 percent less memory per agent.
Supported hardware is deliberately broad: Apple Silicon, NVIDIA, AMD, or a CPU alone. The project says there is no fixed minimum, and that smaller machines run smaller models. It ships as a desktop app with a CLI, connects to Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi and Cline, and exposes an OpenAI-compatible API for everything else. It is Apache 2.0 licensed and runs locally.
The two projects are not competitors. Quail targets the analytic SQL layer, where a query planner generates millions of small, structured model calls. Magnitude targets the developer's laptop, where one agent runs against one model and the constraint is memory and decode speed. But they point at the same cost problem from opposite ends: inference is expensive because general-purpose engines do not know what the workload is going to do next.
What the price data says
AICostBudget's September pricing index, frozen at a 30 September UTC cutoff across 81 public records from 11 providers, gives the market context. Across 69 records with a current input-token scalar, prices range from $0.10 to $30 per 1M tokens. Output prices range from $0.10 to $180. The median input price is $1.32, the median output price is $6, and the median output-to-input ratio is 5.0x.
The index makes a point that is easy to miss in a single cheapest-model ranking. Across 58 comparable records, the median cached-input discount is 90 percent. Across 44 batch-capable records, the median batch discount is 50 percent. A workload comparison that ignores either mode is comparing list prices nobody pays.
Provider medians in the index, which the methodology says are not usage-weighted or market share: Google Gemini at $0.75 input and $4.125 output across 14 priced records, Mistral AI at $0.20 and $0.60 across 7, DeepSeek at $0.30 and $1.20 across 3, OpenAI at $2 and $11 across 18, xAI at $1.25 and $2.50 across 9, Cohere at $1.50 and $5.75 across 2, Anthropic at $5 and $25 across 13, and Moonshot AI at $0.95 and $4 across 3. AWS, Azure and Google Cloud each contribute one record with no numeric value.
The hardware side
Hardware efficiency is being attacked in parallel. EnerTune, a paper published in the SOSP '26 proceedings on 28 September, argues that optimizing GPU clusters purely for utilization can increase energy consumption. The authors, Prasoon Sinha, Dimitrios Liakopoulos, Nathan Lemma and Neeraja J. Yadwadkar, propose analytical models for per-model performance and power, plus the power draw of colocated models on shared GPUs, feeding an energy-aware bin-packing algorithm. They report meeting performance SLOs while cutting energy 1.4x to 2.3x and power draw 1.3x to 2.6x against state-of-the-art baselines.
That is a different lever from Quail's, and it comes with the same caveat: benchmark results, not production deployments. The paper explicitly notes that profiling every model under hundreds of configurations is prohibitively expensive at scale, which is why it uses analytical models instead.
Meanwhile Pinecone's experience with its own pricing calculator, described by Gkogan on 30 September, is a reminder that cost confusion is not only a compute problem. The company removed the calculator after an A/B test found visitors who did not see it were 16 percent more likely to sign up and 90 percent more likely to make contact, with no increase in pricing support tickets. Seven of every ten employees had predicted the calculator version would win.
The common thread across Quail, Magnitude and EnerTune is that the biggest wins come from knowing more about the workload, whether that is a query plan, a specific chip, or a colocation pattern. General-purpose engines leave that information on the table. What none of the three establishes yet is whether the gains survive contact with messy production traffic, and only Magnitude ships something a developer can install today.
Sources
6- 01Hitting 1B tokens/minute on 1 GPU combining a query planner and inference engineEN
- 02Jointly optimizing SQL queries and LLM inference for up to 14x speedupsEN
- 03Launch HN: Magnitude (YC S25) - Self-optimizing inference engine for agentsEN
- 04AI API Pricing Index - September 2026EN
- 05Beyond Utilization: Energy-Conscious GPU Sharing for Inference ServingEN
- 06The pricing calculator made people more confused about pricingEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.