Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Inference Engine Quail Tops 1B Tokens a Minute as Vendor Prices Climb

Quail, an inference engine built by the Full Stack Data Lab, processed more than one billion tokens per minute on a single H100 GPU in results published on 30 September, over 10x faster than its vLLM baseline on the same hardware and at under 6 cents per billion tokens on Modal.

AI & modelsNewsRachel NwosuPublished: 30 September 20267 min readSources 9
Inference Engine Quail Tops 1B Tokens a Minute as Vendor Prices Climb

The numbers land as GPU capacity gets more expensive. AWS has reportedly flagged a 15% hike in GPU capacity prices effective 7 October, and Nebius raised inference prices 18.3% in late September, per separate reports. Against that backdrop, a string of inference-layer projects published benchmarks this week that attack cost from the software side.

Quail, short for QUery-Aware Inference Layer, comes out of the Full Stack Data Lab. The team describes it as a joint optimisation of the SQL query planner and the inference engine, built for AI-SQL: SQL extended with user-defined functions that call an LLM. The engine's own benchmark write-up reports 1.84x faster than vLLM geometrically averaged over tasks, including two queries the authors designed to show where AI-SQL inference still needs work.

The workload looks nothing like a chatbot.

An AI function evaluates its prompt row by row, so a single SQL query can create hundreds of thousands or millions of model calls, according to the lab's technical post, published on 30 September. A filter needs one LLM call per row; a naive join needs one call for every pair of rows across two input tables. Sending those as separate requests to a general-purpose engine such as vLLM carries a large overhead, the post argues, because the requests share prefixes that could be reused.

The lab's BIO-4 query in its QUAIL-B benchmark illustrates the scale: 5,000 long medical reports joined against 4,144 reaction terms, with the reaction list used twice for two joins. The plan filters the term sets for cardiovascular and neurological reactions, then joins the surviving reports against each. Executing it naively would mean a prompt per filter input and a prompt per candidate report-reaction pair.

Quail's answer is to make the planner and the engine aware of each other. Quail is published as version 0.1.0, with an example that runs a query against the IMDB dataset using qwen3-4b-fp8 on an h100-sxm device. The engine's benchmark write-up reports the billion-tokens-per-minute figure on one multi-join query where planning matters most, more than 10x faster than the vLLM baseline on identical hardware.

Modal, which hosts the demo, puts the cost at under 6 cents per billion tokens. That figure applies to a specific query profile: small open-weights models handling short, bounded decisions rather than long agentic chains. The lab's authors name DocETL from UC Berkeley, LOTUS from Stanford, Palimpzest from MIT and ThalamusDB from Cornell among the academic systems in the same field, and credit techniques such as MOAR, Task Cascades, Abacus and BARGAIN for cutting call counts and model sizes.

Their claim is that even after those optimisations, a plan can still require hundreds of thousands or millions of calls, which is where a purpose-built engine earns its keep.

Quail is not alone in targeting local hardware. Magnitude, a YC S25 company, launched an open source inference engine for agents on 30 September, publishing it on GitHub under Apache 2.0. Its README claims open models run up to 2x faster than llama.cpp, broken down as 92% faster decode on Metal and 19% on CUDA. The mechanism is compilation and kernel tuning on the user's own device before a model runs, rather than shipping precompiled kernels for broad hardware classes.

Magnitude also claims 27% less memory per agent, freed when agents stop, and prefix cache sharing across concurrent sessions. It supports Apple Silicon, NVIDIA, AMD or CPU-only, and connects to Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi and Cline, with an OpenAI-compatible API for anything else. The company says prompts, files and models stay on the machine, with no internet needed once a model is downloaded.

Both projects are pitching the same economics: move inference off metered cloud endpoints and onto hardware the buyer already owns or rents at a flat rate. The pricing data suggests why that pitch is landing.

AICostBudget's September pricing index, frozen at a 30 September cutoff, covers 81 public records across 11 providers. Median input price across 69 non-null scalars is $1.32 per million tokens; median output price is $6. Median output runs 5.0x input. Observed ranges are wide: input spans $0.10 to $30 per million tokens, output $0.10 to $180.

The index also quantifies the discount structures that make headline rates misleading. Across 58 comparable records, the median cached-input discount is 90%, with an observed range of 75% to 98%. Across 44 batch-capable records, the median Batch input discount is 50%, observed range 20% to 50%. The report's own conclusion is blunt: a workload comparison that ignores cache and Batch pricing omits published savings modes for a large share of tracked records.

Per-provider medians in the same dataset put OpenAI at $2 input and $11 output across 18 records, Google Gemini at $0.75 and $4.125 across 14, Anthropic at $5 and $25 across 13, xAI at $1.25 and $2.5 across nine, Mistral AI at $0.20 and $0.60 across seven, DeepSeek at $0.30 and $1.2 across three, Moonshot AI at $0.95 and $4 across three, and Cohere at $1.5 and $5.75 across two. The report cautions that document pages, storage, session duration and other time-based units cannot be flattened into a token-only table, and that 13 records carry non-token components.

Not every cost problem is a pricing problem. A paper presented at SOSP '26 on 28 September argues that optimising GPU clusters for utilisation alone can raise energy consumption. The authors, Prasoon Sinha, Dimitrios Liakopoulos, Nathan Lemma and Neeraja J. Yadwadkar, present EnerTune, which uses analytical models of per-model performance and power, plus the power draw of colocated models on shared GPUs, inside an energy-aware bin-packing algorithm that picks placement and configuration together.

Their abstract reports energy reductions of 1.4x to 2.3x and power draw reductions of 1.3x to 2.6x over state-of-the-art baselines while meeting performance SLOs. The paper argues that profiling every model across hundreds of configurations is too expensive at scale, which is why the approach relies on models rather than measurement.

Latency, not throughput, is the constraint in physical AI. MindOn published Mind-1 on 30 September, claiming it cuts robot inference latency from 82ms to 32ms. The company's post notes that at an end-effector speed of 3 m/s, 10ms of latency corresponds to roughly 3cm of movement, enough to affect precise grasping, insertion or contact-rich manipulation. MindOn also points at training data tempo: much manipulation data comes from teleoperation, which is slower than natural human motion, and models learn the tempo along with the task.

The same efficiency logic shows up outside AI. CleanTechnica reported on 29 September that Transport & Environment analysis puts EV running costs at less than half of diesel per kilometre, based on 6.9 L/100 km diesel consumption against 20.2 kWh/100 km electric and an electricity price of €0.343/kWh. T&E's Antony Froggatt said the EU should tax oil profits to fund support for low-income households and scrappage schemes. T&E's charging estimate is described in its own methodology as likely an upper bound.

Telecoms operators are pursuing the same playbook through consolidation. Light Reading reported on 30 September that StarHub and M1 are in merger talks in Singapore, where the pair already share 5G spectrum and radio access network through a jointly owned company called Antina. Singtel reportedly held a 43% mobile share in June, with M1 at 22%, StarHub at 21% and Simba at 14%. In New Zealand, 2degrees and OneNZ have proposed combining their radio networks into a jointly owned entity, with a decision from the Commerce Commission expected next year. Light Reading cites the China Telecom and China Unicom 5G arrangement as the largest such partnership, covering 1.5 million basestations and claimed capex savings of $56 billion.

Whether Quail's throughput translates into lower bills for ordinary AI-SQL users is not yet established. The 1B tokens per minute figure comes from one multi-join query chosen because planning matters most, and the engine's averaged benchmark over the full task set is 1.84x, not 10x. The lab itself included two queries designed to expose current weaknesses. Quail is at version 0.1.0.

Comments 0

Sources

9
  1. 01Jointly optimizing SQL queries and LLM inference for up to 14x speedupsEN
  2. 02Hitting 1B tokens/minute on 1 GPU combining a query planner and inference engineEN
  3. 03AI API pricing index - September 2026EN
  4. 04Launch HN: Magnitude (YC S25) - Self-optimizing inference engine for agentsEN
  5. 05Beyond Utilization: Energy-Conscious GPU Sharing for Inference ServingEN
  6. 06Mind-1: Cutting robot inference latency from 82ms to 32msEN
  7. 07The pricing calculator made people more confused about pricingEN
  8. 08Driving An EV Now Costs Half As Much As Diesel, New Analysis ShowsEN
  9. 09Singapore, New Zealand operators seek partnerships to cut costsEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.