Quail Claims 1B Tokens Per Minute On A Single H100 As Inference Costs Split
Two new inference engines claim order-of-magnitude efficiency gains on the same GPU hardware, while OpenAI's new $500 subscription and a September pricing snapshot show the per-token market is still moving in opposite directions.

On 30 September, the Full Stack Data Lab published a blog describing Quail, an inference engine built specifically for AI-SQL, the SQL extensions that let queries call large language models row by row. The project claims more than a billion tokens processed per minute on a single H100 GPU on one multi-join query, and a 1.84x geometric-mean speedup over vLLM across a new benchmark of AI-SQL workloads.
That is the headline number. It is also a narrow one.
The billion-tokens-per-minute figure comes from a single query the team describes as one where "planning is particularly important," not from an average. The broader benchmark number, 1.84x, is the one the authors present as the general result. The same day, Magnitude, a Y Combinator S25 company, launched an open source inference engine it says tunes its kernels on the user's own hardware before a model runs. Its GitHub README claims up to 2x faster than llama.cpp, with 92% faster decode on Apple Metal and 19% on CUDA, and 27% less memory per agent. The project ships under Apache 2.0 and runs on Apple Silicon, NVIDIA, AMD, or CPU only.
What the numbers actually compare
Neither claim is a general statement about the price of inference. Magnitude benchmarks against llama.cpp, a general-purpose engine, and attributes its edge to hand-written kernels for popular open-weight model families. Quail benchmarks against vLLM, and attributes its edge to joint optimization of the SQL query plan and the inference engine, so that millions of related model calls share prefix caches instead of arriving as separate requests.
The Full Stack Data Lab post, dated 24 September, lays out the cost problem directly: an AI function in a SQL query evaluates its prompt per row, so a filter needs one LLM call per row and a naive join needs one call for every pair of rows in its two inputs.
The team's benchmark query, BIO-4, runs over 5,000 long medical reports joined against 4,144 reaction terms. The authors say sending those calls to a general-purpose engine as separate requests carries a large cost. Their fix is to plan at the query level. Modal's blog on the same work puts the economics more bluntly: under 6 cents per billion tokens on Modal hardware for the billion-TPM query. That figure is a vendor's own number for its own platform, and it covers one query shape, not a representative workload.
Different inference applications produce different inference workloads, and AI-SQL is no exception. A query like the one above might produce millions of sequences of thousands of tokens.
Subscription pricing moves the other way
On the consumer side, OpenAI went the other direction. At its Tuesday DevDay, the company announced a Pro 500 tier at $500 a month, $300 more than its previous top plan, according to Business Insider. The tier includes access to "Ultrafast" compute for GPT-6 Astra in ChatGPT Work and Codex, which OpenAI says is an 8x speed increase in Codex, and the company's highest included usage.
OpenAI also reopened its $200 Pro plan but halved the compute allowance.
The allowance is now 10x the $20 Plus plan, down from a 20x multiple, and GPT-6 Pro chat messages drop from 200 to 100 a week. Thibault Sottiaux, engineering lead for Codex, wrote in a Monday X post that the changes net out at half the dollar in API spend versus the prior plan. "I wanted to make sure to share this change ahead of time so you can all understand it before we shower you with good news," he wrote, as quoted by Business Insider. The two moves point in opposite directions: a cheaper-per-token allowance at the $200 level, a much more expensive tier on top. Existing Pro subscribers get a one-time credit during the transition.
A separate snapshot published 30 September by AICostBudget covers 81 public pricing records across 11 providers and puts median input at $1.32 and median output at $6 per million tokens. The observed range across 69 records with an input scalar is $0.10 to $30, and $0.10 to $180 for output. Output is commonly the higher list-price component, with a median output-to-input ratio of 5.0x.
The same dataset reports a median cached-input discount of 90% across 58 comparable records and a median Batch input discount of 50% across 44. That is the part of the market where the effective price can diverge sharply from the list price, and it is the part a naive cost comparison usually misses.
Hardware, energy and the rest of the stack
Efficiency work is not only about tokens per second. A paper presented at SOSP '26 on 28 September describes EnerTune, a serving system that models per-model power draw and the draw of models colocated on shared GPUs, then uses an energy-aware bin-packing algorithm to place and configure them. The authors report 1.4x to 2.3x lower energy consumption and 1.3x to 2.6x lower power draw than state-of-the-art baselines, while meeting performance SLOs. Their argument is that optimizing purely for utilization can raise energy use, which matters when the bill is not only per token.
On the robotics side, MindOn introduced Mind-1 on 30 September, a physical AI model it says cuts robot inference latency from 82ms to 32ms.
The company notes that at an end-effector speed of 3 m/s, 10ms of latency corresponds to roughly 3cm of movement, enough to affect precise grasping. That is a different market from data-center inference, but the same constraint: latency and cost are set by the whole pipeline, not the model alone.
Telecoms operators are pursuing a related logic on infrastructure costs. Light Reading reported on 30 September that StarHub and M1 are in merger talks in Singapore, where they already share 5G spectrum and radio access network through a joint company called Antina. In New Zealand, 2degrees and OneNZ have proposed pooling their radio networks into a jointly owned entity. The article cites Singtel's reported 43% share of Singapore's mobile market in June, against 22% for M1, 21% for StarHub and 14% for Simba, and says a combined StarHub-M1 would be on a similar scale to Singtel. The China Telecom and China Unicom 5G arrangement is cited as the largest such sharing deal, covering 1.5 million basestations and claimed savings of $56 billion in capex.
Not every attempt to cut cost by adding machinery works. A post published 30 September by Gokul Kogan describes how Pinecone removed a usage-based pricing calculator after A/B testing. Visitors who did not see the calculator were 16% more likely to sign up and 90% more likely to contact the company, with no increase in pricing support tickets, according to the post. In an internal poll, seven of ten employees expected the version with the calculator to perform better. The lesson is a caution for anyone building a cost estimator for a token-priced service: an estimate that is wrong by an order of magnitude can lose the customer entirely.
Put together, the picture is less a single falling price than a set of engineering bets. Quail bets on query-aware batching. Magnitude bets on per-device kernel compilation. EnerTune bets on power-aware placement. OpenAI is testing how much a small number of heavy users will pay for speed, while its cheaper tier gets less compute than before. The only number that is stable across the dossier is the shape of the bill: output tokens cost more than input, cache and batch discounts are large, and the headline figures from any single benchmark describe one query, on one machine, on one day.
Sources
10- 01Jointly optimizing SQL queries and LLM inference for up to 14x speedupsEN
- 02Hitting 1B tokens/minute on 1 GPU combining a query planner and inference engineEN
- 03Launch HN: Magnitude (YC S25) - Self-optimizing inference engine for agentsEN
- 04ChatGPT's new plan costs $500 a monthEN
- 05AI API pricing index - September 2026EN
- 06Beyond Utilization: Energy-Conscious GPU Sharing for Inference ServingEN
- 07Mind-1: Cutting robot inference latency from 82ms to 32msEN
- 08Singapore, New Zealand operators seek partnerships to cut costsEN
- 09The pricing calculator made people more confused about pricingEN
- 10Driving An EV Now Costs Half As Much As Diesel, New Analysis ShowsEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.