Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Quail Claims 1B Tokens a Minute on One GPU as Inference Cost Debate Sharpens

A research group says its query-aware inference layer processed more than a billion tokens per minute on a single Nvidia H100 GPU, over 10x faster than a vLLM baseline, in the newest push to cut the cost of running large language models at scale. The claim was published on 30 September, the same day OpenAI announced a $500-a-month ChatGPT tier and an open source engine promised 2x faster local inference.

AI & modelsExplainerGrace OkonkwoPublished: 30 September 20265 min readSources 10
Quail Claims 1B Tokens a Minute on One GPU as Inference Cost Debate Sharpens

The numbers come from the Full Stack Data Lab and Modal, which published separate posts on 30 September about a system called Quail, short for QUery-Aware Inference Layer. Modal's blog states that on one multi-join query, Quail hit more than a billion tokens processed per minute per H100 GPU. That, it says, is over 10x faster than its vLLM baseline on the same hardware. On Modal, it works out to under 6 cents per billion tokens.

The two posts describe the same project from different angles. The Full Stack Data Lab post, dated 24 September but published on the site on 30 September, explains the problem. AI-SQL extends SQL with functions that call models row by row, so one filter can mean one model call per row and a join can mean one call for every pair of rows. The lab's benchmark query, BIO-4 in a suite called QUAIL-B, runs over 5,000 medical reports and 4,144 reaction terms.

Small models, huge volumes

Modal's write-up frames AI-SQL as a quieter sibling of the agent and chatbot boom. Its example query joins customers and products with a prompt asking whether one might buy the other. Such a query, Modal writes, can produce millions of sequences of thousands of tokens, and those sequences often need less than frontier intelligence. Small open-weight models can handle them. The bottleneck is not model quality, according to Modal. It is the cost of delivering millions of related requests to an engine built for arbitrary user-controlled traffic.

Quail's approach is to plan the SQL query and the inference workload together rather than treating the engine as a generic endpoint. Modal says the geometric mean speedup over vLLM across its new AI-SQL benchmark is 1.84x. That is a much smaller figure than the headline billion-tokens-per-minute case. The company is explicit that two queries in the benchmark were designed to expose areas for future improvement.

"I see it as a point on the LLM pareto optimal curve in a regime that had a large revealed latent demand (no thinking, single token, low latency acceptable intelligence) that was under-invested into because of a race to higher intelligence."

Modal attributes that quote to Karpathy, describing the Jev model from TypeSafe AI. The same post notes that many database vendors now offer AI functions. The list includes Snowflake Cortex AISQL, BigQuery AI functions, Databricks AI Functions and, recently, MotherDuck.

Academic systems have been attacking the same waste. The Full Stack Data Lab lists DocETL from UC Berkeley, LOTUS from Stanford, Palimpzest from MIT and ThalamusDB from Cornell. It also names techniques such as MOAR, Task Cascades, Abacus and BARGAIN that cut the number of calls or pick cheaper models. Even then, the lab writes, a plan may still need hundreds of thousands or millions of calls.

Price list versus price structure

While Quail attacks the cost of a query, an independent snapshot published on 30 September tracks what providers charge. The AICostBudget AI API Pricing Index covers 81 frozen public records across 11 providers at a 2026-09-30 UTC cutoff. Across 69 records with an input-token scalar, the observed range runs from $0.10 to $30 per million tokens. Output prices span $0.10 to $180. The median input price is $1.32 and the median output price is $6, a ratio of 5.0x.

The index makes a point that matters for anyone modelling inference spend: cache and batch pricing are not rounding errors. Across 58 comparable records, the median cached-input discount is 90%. Across 44 batch-capable records the median batch discount is 50%. A workload comparison that ignores both, the report says, omits published savings modes for a large share of tracked records. Provider-level medians in the snapshot include OpenAI at $2 input and $11 output across 18 records, Google Gemini at $0.75 and $4.125 across 14, and Anthropic at $5 and $25 across 13. Mistral AI, DeepSeek, Cohere and Moonshot AI also appear.

End-user pricing moved in the opposite direction on the same day. Business Insider reported that OpenAI announced a new Pro 500 tier at its Tuesday DevDay, costing $500 a month, a $300 increase over the previous top plan. It comes with access to "Ultrafast" computing for GPT-6 Astra in ChatGPT Work and Codex and an 8x speed claim for Codex. OpenAI also reopened its $200 Pro plan while halving its compute allowance to 10x the $20 Plus plan, down from 20x. It cut GPT-6 Pro chat messages from 200 to 100 a week. Thibault Sottiaux, engineering lead for Codex, said in a Monday X post that the changes net out at half the dollar in API spend compared with the prior plan.

On the hardware side, the picture is messier. A GitHub launch post for Magnitude, a YC S25 company, claims its open source engine tunes kernels on the user's device and runs open models up to 2x faster than llama.cpp. It cites 92% faster decode on Metal and 19% on CUDA, plus 27% less memory per agent. Those are vendor benchmarks, not independent tests, and the post gives no third-party verification.

Separately, the SOSP '26 paper "Beyond Utilization" from Prasoon Sinha, Dimitrios Liakopoulos, Nathan Lemma and Neeraja J. Yadwadkar argues that optimizing GPU clusters purely for utilization can raise energy consumption. It presents a system called EnerTune that it says cuts energy by 1.4-2.3x and power draw by 1.3-2.6x over state-of-the-art baselines. The paper was published on 28 September. On 30 September, CleanTechnica reported analysis by Transport & Environment finding that EV drivers in Europe pay less than half what diesel drivers pay per kilometre, with diesel €32 more per tank than at the start of the year and about €16 of that extra refinery margin.

Comments 0

Sources

10
  1. 01Jointly optimizing SQL queries and LLM inference for up to 14x speedupsEN
  2. 02Hitting 1B tokens/minute on 1 GPU combining a query planner and inference engineEN
  3. 03AI API Pricing Index - September 2026EN
  4. 04ChatGPT's new plan costs $500 a monthEN
  5. 05Launch HN: Magnitude (YC S25) - Self-optimizing inference engine for agentsEN
  6. 06Beyond Utilization: Energy-Conscious GPU Sharing for Inference ServingEN
  7. 07Driving An EV Now Costs Half As Much As Diesel, New Analysis ShowsEN
  8. 08Mind-1: Cutting robot inference latency from 82ms to 32msEN
  9. 09The pricing calculator made people more confused about pricingEN
  10. 10Singapore, New Zealand operators seek partnerships to cut costsEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.