Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Quail Runs AI-SQL at 1B Tokens a Minute on One H100, Beating vLLM 10x

Modal and Carnegie Mellon's Full Stack Data Lab published Quail on 30 September. The inference engine plans SQL queries and schedules LLM execution together, and on a single H100 it clears one billion tokens per minute, more than 10x a vLLM baseline.

AI & modelsAnalysisGrace OkonkwoPublished: 30 September 20263 min readSources 7
Quail Runs AI-SQL at 1B Tokens a Minute on One H100, Beating vLLM 10x

The numbers are the story. On one multi-join query, Quail pushes past a billion tokens per minute on a single H100 GPU, according to Modal, which built the engine with Carnegie Mellon University's Full Stack Data Lab. On Modal's own hardware that works out to under 6 cents per billion tokens. On a new benchmark for AI-SQL queries, Quail runs 1.84x faster than vLLM, geometrically averaged across tasks.

A companion post from the Full Stack Data Lab, published the same day, shows where the baseline breaks down. The team ran vLLM 0.26.0 with Qwen3 4B FP8 on one H100 against BIO-4, a query that filters 5,000 medical reports against 4,144 reaction terms and joins them twice. Their speed-of-light estimate for the plan was 894.37 seconds, or 14.91 minutes. vLLM took 6.84 hours, 27.55x that estimate.

Two causes. Host overhead, where the CPU schedules and tracks requests while the H100 idles. And KV regret: vLLM throws away key-value cache it needs again later, so at scale factor 1.0 the run processed 174.6 million tokens of wasted prefill.

Quail's fix is structural. With a query plan in hand, the engine can order and evict KV cache deliberately, a revision of Hydragen-style cascade attention, and it can batch small requests so the GPU stays busy. Modal describes the result as the point where database planning and inference scheduling stop being separate problems.

Why this matters for the bill

AI-SQL is the pattern where SQL is extended with user-defined functions that call an LLM. Snowflake Cortex AISQL, BigQuery AI functions, Databricks AI Functions and, recently, MotherDuck all support it, the Full Stack Data Lab notes. A filter means one model call per row. A naive join means one call for every pair of rows. One query can generate hundreds of thousands or millions of calls.

Against that backdrop, the general price picture is getting messier, not cleaner. AICostBudget's September pricing index, frozen at the 2026-09-30 UTC cutoff, covers 81 public records across 11 providers. Median input price is $1.32 per million tokens; median output price is $6. The median cached-input discount is 90%, and the median batch discount is 50%. The same page warns that list prices are planning inputs, not invoices.

Raw hourly GPU rates tell a similar story of divergence. GPUAdvisor, which tracks 14 providers, last tracked on 1 October, lists an A4000 16GB at Hyperstack for $0.15 per hour, an RTX 3090 at Lium for $0.22, and an L40S at Lium for $0.38, which it marks as 89% below AWS. The site's own summary is blunt: GPU-specialist clouds offer 2 to 4 times lower per-GPU pricing than hyperscalers, and spot can save 60 to 70% on training runs with checkpointing.

FitMyLLM converts hourly rates into cost per million tokens, and the ranking inverts. Its cheapest measured option is an RTX 3060 12GB on Vast.ai at $0.261 per million tokens. An RTX 5090 at Vast.ai is the fastest measured card at 227 tokens per second but costs $0.576 per million. Datacenter cards rank badly on this single-stream view, and FitMyLLM says so: an H100 earns its price by serving many streams at once, and at one stream its bandwidth sits mostly idle.

The efficiency question underneath

Energy is the next axis. A paper presented at SOSP '26 on 28 September, from Prasoon Sinha, Dimitrios Liakopoulos, Nathan Lemma and Neeraja J. Yadwadkar, argues that optimising purely for GPU utilisation can raise energy consumption. Their system, EnerTune, claims 1.4 to 2.3 times lower energy and 1.3 to 2.6 times lower power draw over state-of-the-art baselines while meeting performance SLOs.

Not every vendor is chasing the same curve. MindOn's Mind-1, announced 30 September, targets robot inference latency, cutting it from 82ms to 32ms, a different corner of the same problem: what an hour of compute actually buys.

Comments 0

Sources

7
  1. 01Hitting 1B tokens/minute on 1 GPU combining a query planner and inference engineEN
  2. 02Jointly optimizing SQL queries and LLM inference for up to 14x speedupsEN
  3. 03AI API pricing index - September 2026EN
  4. 04Cloud GPU Pricing - Cost Intelligence for H100, A100 & B200EN
  5. 05Cloud GPU prices for LLM inference - cost per million tokensEN
  6. 06Beyond Utilization: Energy-Conscious GPU Sharing for Inference ServingEN
  7. 07Mind-1: Cutting robot inference latency from 82ms to 32msEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.