Open source inference engines cut AI-SQL and agent costs, benchmarks show
Two open source inference projects published benchmarks on 30 September claiming large cuts in the cost of running AI workloads, as a separate pricing index put the median output token at $6 per million.

On 30 September, Modal and Carnegie Mellon University's Full Stack Data Lab launched Quail. It is an inference engine that optimises SQL query planning and LLM execution together. Modal's blog says Quail processes more than a billion tokens per minute on a single H100 GPU. That is over 10x faster than its vLLM baseline on the same hardware. On Modal's infrastructure, the blog says, that works out to under 6 cents per billion tokens.
The Full Stack Data Lab blog, published on 24 September and surfaced alongside the Modal post, explains the problem. AI-SQL extends SQL with functions that call an LLM per row, so a single filter query can trigger hundreds of thousands or millions of model calls. The lab attributes support to Snowflake Cortex AISQL, BigQuery AI functions, Databricks AI Functions and, recently, MotherDuck.
Quail is the newest of several efficiency claims this week. On 30 September, the open source project Magnitude launched on GitHub. It describes an inference engine for agents that compiles and tunes kernels on the user's own hardware. Its README claims decode speeds up to 2x faster than llama.cpp, with 92% faster decode on Metal and 19% on CUDA, and 27% less memory per agent. Magnitude runs on Apple Silicon, NVIDIA, AMD or CPU only. It connects to Pi, OpenCode, Hermes and Codex through a single click.
On one multi-join query where planning is particularly important, Quail hits over a billion tokens processed per minute per H100 GPU (TPM/GPU), >10x faster than our vLLM baseline on the same hardware.
The two projects attack different bottlenecks. Magnitude targets local agent workloads, where the constraint is kernel fit to consumer or workstation chips. Quail targets database workloads, where the constraint is host overhead and KV cache eviction. The Full Stack Data Lab measured a vLLM baseline on BIO-4, a medical-report query, at 6.84 hours. That is 27.55x its own speed-of-light estimate of 14.91 minutes.
Hardware economics remain the backdrop. A paper presented at SOSP '26 and published on 30 September describes EnerTune, a GPU sharing system. The authors say it cuts energy consumption by 1.4 to 2.3 times and power draw by 1.3 to 2.6 times over state-of-the-art baselines while meeting performance SLOs. The authors argue that optimising purely for GPU utilisation can raise energy consumption. They also say that profiling every configuration at scale is prohibitively expensive.
List prices stay wide
On the pricing side, AICostBudget published its September index on 30 September. It is based on 81 frozen public pricing records across 11 providers. It puts the median input price at $1.32 per million tokens and the median output price at $6, a ratio of 5.0x. The observed range runs from $0.10 to $30 per million for input and $0.10 to $180 for output.
The index reports a median cached-input discount of 90% across 58 comparable records. It reports a median Batch discount of 50% across 44 records. By provider median input price, Mistral AI sits lowest at $0.20. DeepSeek is at $0.30, Google Gemini at $0.75, Moonshot AI at $0.95, xAI at $1.25, Cohere at $1.50, OpenAI at $2 and Anthropic at $5.
Not every record is directly comparable. AICostBudget says 17 records use tiered or long-context pricing and 13 use non-token components such as document pages or session duration. These are excluded from its token medians. It also notes that its snapshot window begins in July 2026 and does not claim an industry-wide annual price decline.
Meanwhile, the engineering cost of not optimising is visible in the Quail numbers. The lab says vLLM processed 174.6 million tokens it did not need on BIO-4, due to KV regret, and that CPU scheduling left the H100 idle. Those are the gaps Magnitude and Quail are both trying to close, from opposite ends of the stack.
Sources
5- 01Hitting 1B tokens/minute on 1 GPU combining a query planner and inference engineEN
- 02Jointly optimizing SQL queries and LLM inference for up to 14x speedupsEN
- 03Launch HN: Magnitude (YC S25) - Self-optimizing inference engine for agentsEN
- 04AI API pricing index - September 2026EN
- 05Beyond Utilization: Energy-Conscious GPU Sharing for Inference ServingEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.