Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Quail Claims 1B Tokens a Minute on One H100 as GPU Price Tracking Gets Crowded

Modal and Carnegie Mellon's Full Stack Data Lab say their Quail engine pushed a multi-join AI-SQL query past one billion tokens per minute on a single H100, more than 10x their vLLM baseline, in a post published on 30 September.

AI & modelsAnalysisRachel NwosuPublished: 30 September 20265 min readSources 7
Quail Claims 1B Tokens a Minute on One H100 as GPU Price Tracking Gets Crowded

The claim comes with a price tag attached: under 6 cents per billion tokens on Modal, according to the company's blog. That is the sort of number that decides whether AI-SQL gets deployed or quietly shelved. A filter over a million rows is a million model calls.

Two bottlenecks, two fixes

The Full Stack Data Lab's own write-up, published 24 September, frames the problem in database terms. A query joining 5,000 medical reports against 4,144 reaction terms expands into millions of separate inference requests. Running them one at a time through vLLM 0.26.0 with Qwen3 4B FP8 on one H100 took 6.84 hours. The authors calculate that is 27.55x their roofline estimate of 894.37 seconds, or 14.91 minutes. They name two causes. The first is host overhead, where the CPU schedules requests while the H100 idles. The second is KV regret, where the engine discards key-value state it later needs again. On that benchmark at scale factor 1.0, vLLM processed 174.6 million tokens it did not have to.

Quail's answer is to look at the query plan before running it. If you know the shape of the joins, you can order requests so shared prefixes stay cached. You can also batch work that a general-purpose engine treats as unrelated. The Modal post describes it as a revision of Hydragen-style cascade attention. Both teams call the project a collaboration between Modal's inference researchers and CMU's database group.

Over the full benchmark, the number is smaller: 1.84x faster than vLLM, geometrically averaged. That figure includes two queries the authors built specifically to show where AI-SQL inference still fails.

The pricing data underneath

Cost claims land differently depending on what a GPU hour costs in the first place, and that number is now moving weekly. GPUAdvisor tracks 14 providers. It lists Hyperstack at $0.15 per hour for an A4000 16GB, Lium at $0.22 for an RTX 3090 and $0.38 for an L40S, and RunPod's RTX A5000 at $0.27 on-demand against $0.16 on its shared Community Cloud tier. Its page, last tracked 1 October, states plainly that GPU pricing stopped being a stable lookup table.

FitMyLLM attacks the same question from the token side. An hourly rate tells you nothing without a throughput figure, the site argues. Its live table reads rates from provider APIs every three minutes and ranks by cost per million output tokens at one stream. Vast.ai's RTX 3060 listing comes out cheapest at $0.261 per million tokens. An RTX 5090 on Vast.ai runs at 227 tokens per second and $0.576 per million. The site is upfront that datacenter cards rank badly on this metric by design, since an H100 or H200 earns its price by serving many streams at once. It also warns that those rows are estimates rather than measurements.

"A marketplace price is one host's listing, not a published rate: it can vanish, and hosts may run a card below its stock power limit, which costs speed the price does not show."

For API buyers rather than GPU renters, AICostBudget's September index freezes 81 public records across 11 providers at a 2026-09-30 cutoff. Median input price is $1.32 per million tokens, median output $6, and the median output-to-input ratio is 5.0x. Cached-input records show a median 90% discount and Batch-priced records a median 50%, across 58 and 44 comparable rows respectively. The index explicitly warns that its provider medians are not usage-weighted market prices. It also says it does not claim an industry-wide annual decline.

Efficiency moves up the stack

Energy is the next line item. A paper published at SOSP '26 on 28 September by Prasoon Sinha, Dimitrios Liakopoulos, Nathan Lemma and Neeraja J. Yadwadkar argues that optimizing purely for GPU utilization can raise energy consumption. Their system, EnerTune, uses analytical models of per-model performance and power, including the draw of colocated models on a shared GPU, inside an energy-aware bin-packing algorithm. The authors report meeting performance SLOs while cutting energy 1.4x to 2.3x and power draw 1.3x to 2.6x against state-of-the-art baselines. Profiling every configuration is itself too expensive at scale, which is why the models are analytical rather than measured.

On-device inference is making the same argument from the opposite end. Magnitude, a YC S25 company that posted to Hacker News on 30 September, ships an open source engine under Apache 2.0 that compiles and tunes kernels on the user's own hardware before a model runs. Its README claims up to 2x faster than llama.cpp, with 92% faster decode on Metal and 19% on CUDA, and 27% less memory per agent. No token costs, nothing leaves the machine. The trade-off is that it runs open models on your GPU, not a frontier model on someone else's.

What buyers should watch

Three numbers in the dossier point in different directions. Quail's 1.84x average speedup is real but modest next to the billion-tokens-per-minute headline, and the authors say so themselves. FitMyLLM's cheapest measured rate of $0.261 per million tokens sits on a consumer card that most production teams would not deploy. AICostBudget's median output price of $6 per million tokens reflects list prices, not the cached or batched rates most large workloads actually pay.

The practical reading: per-GPU hourly rates are now a commodity comparison, and the differentiation has moved to throughput per dollar, cache behaviour and energy. Quail's authors are explicit that this is a start, not a finish. They are asking for the one thing their benchmark cannot supply on its own, which is real workloads from teams running AI-SQL at scale.

Comments 0

Sources

7
  1. 01Quail: Speeding up AI-SQL by jointly optimizing query planner and inference engineEN
  2. 02Building an Ultra-High Throughput AI-SQL EngineEN
  3. 03Cloud GPU Pricing — Cost Intelligence for H100, A100 & B200EN
  4. 04Cloud GPU prices for LLM inference — cost per million tokensEN
  5. 05AI API Pricing Index — September 2026EN
  6. 06Beyond Utilization: Energy-Conscious GPU Sharing for Inference ServingEN
  7. 07Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agentsEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.