Inference pricing splinters: open engines squeeze tokens, hardware bills climb
An open source inference engine that tunes its CPU kernels on your own device launched on 30 September, claiming decode speeds up to 92% faster than llama.cpp on Metal, as new research and pricing data point to a widening split in the cost of running AI.

On 30 September, Magnitude, a Y Combinator S25 company, published an open source inference engine for agents that compiles and tunes its kernels on the user's own hardware. The repository says open models run "up to 2x faster than llama.cpp," with 92% faster decode on Metal and 19% on CUDA, and 27% less memory per agent. It runs on Apple Silicon, NVIDIA, AMD or a CPU, and connects to Pi, OpenCode, Hermes, Codex, Claude Code and Cline, according to the project's GitHub page.
That is the newest entry in a week of competing claims about what inference should cost. On the same day, researchers behind Quail, a query-aware inference layer, reported running a multi-join AI-SQL query at more than one billion tokens per minute on a single H100 GPU, which they say is over 10x faster than their vLLM baseline on the same hardware.
The token side: planning beats brute force
The Quail work, published on the Full Stack Data Lab blog and detailed further on Modal's blog, attacks a specific workload: AI-SQL, where SQL queries call language models row by row. A filter needs one call per row, and a naive join needs a call for every pair of rows. On a benchmark query the researchers call BIO-4, inputs contain 5,000 long medical reports and 4,144 reaction terms, used twice across two joins. Their argument is that sending millions of related calls to a general-purpose engine as separate requests wastes the shared prefix state. By jointly optimizing the query plan and the inference engine, they cut that overhead. Modal's write-up puts the cost at under 6 cents per billion tokens on its platform, and reports Quail at 1.84x faster than vLLM geometrically averaged across a new AI-SQL benchmark, including two queries the team designed to expose remaining weaknesses.
"Different inference applications produce different inference workloads, and AI-SQL is no exception," the Modal post says, noting that a single query can produce millions of sequences of thousands of tokens.
Both numbers come from the teams building the systems, and neither has been independently reproduced from the published material. That caveat matters, because the wider pricing picture is messier than any single speedup suggests.
AICostBudget's September pricing index, frozen at a 2026-09-30 UTC cutoff, tracks 81 public records across 11 providers. It puts the median input price at $1.32 per 1M tokens and the median output price at $6, a ratio of 5.0x across the 69 records carrying both scalars. Observed ranges run from $0.10 to $30 for input and $0.10 to $180 for output. The same dataset reports a median cached-input discount of 90% across 58 comparable records and a median batch discount of 50% across 44 batch-capable rows. In other words, list prices are a poor guide to what a workload actually pays. Provider medians in the index range from Mistral AI at $0.20 input and $0.60 output, over three records, to Anthropic at $5 and $25 across 13. OpenAI's 18 records sit at $2 and $11; Google Gemini's 14 at $0.75 and $4.125.
The hardware side: energy, not just utilization
If token economics are improving, GPU economics are not obviously following. A paper presented at SOSP '26 on 28 September argues that optimizing GPU clusters purely for utilization can raise energy consumption. The authors, Prasoon Sinha, Dimitrios Liakopoulos, Nathan Lemma and Neeraja J. Yadwadkar, describe EnerTune, a serving system that models per-model performance and power draw, including colocated models on shared GPUs, and uses an energy-aware bin-packing algorithm to place and configure models. The paper claims energy reductions of 1.4-2.3x and power-draw reductions of 1.3-2.6x over state-of-the-art baselines while meeting performance SLOs. It also makes a practical point about scale: profiling every model under hundreds of configurations is prohibitively expensive, and the profiling itself burns energy.
That is an academic result, not a vendor benchmark, and it lands as operators in Asia-Pacific take a different route to cutting costs. Light Reading reported on 30 September that StarHub and M1 are in merger talks in Singapore, where they already share 5G spectrum and radio access network through a jointly owned company called Antina. Singtel reportedly held a 43% share of the mobile market in June, with M1 at 22%, StarHub at 21% and Simba at 14%.
In New Zealand, 2degrees and OneNZ have proposed combining their radio networks into a jointly owned entity, a deal expected to reach the Commerce Commission next year. Light Reading notes the precedent: China Telecom and China Unicom's shared 5G infrastructure now amounts to 1.5 million basestations and claimed savings of $56 billion in capex. These are connectivity deals rather than AI compute deals, but they show the same instinct, that sharing expensive fixed assets beats duplicating them.
What the buyers see
For anyone actually paying for inference, the most useful data point of the week may be a cautionary one. Gaurav Kogan, writing on 30 September, described removing a usage-based pricing calculator at Pinecone after an A/B test. Visitors who did not see the calculator were 16% more likely to sign up and 90% more likely to contact the company, he wrote, with no increase in pricing support tickets. His account of why the calculator was removed is specific: a slight misinterpretation could produce an estimate overstated by as much as 1,000x, and the calculator gave users false confidence. An internal poll found 7 of every 10 employees expected the version with the calculator to perform better. The lesson transfers to inference pricing, where cached and batch rates can move a bill by an order of magnitude and the headline number rarely reflects the workload.
Two other 30 September items sit at the edge of this story. MindOn introduced Mind-1, a physical AI model it says cuts robot inference latency from 82ms to 32ms, and notes that at an end-effector speed of 3 m/s, 10ms of latency equals roughly 3cm of movement. And Transport & Environment published analysis, covered by CleanTechnica on 29 September, claiming EV drivers pay less than half what diesel drivers pay per kilometre, with diesel drivers paying 32 euros more per tank than at the start of the year.
Neither is a data-centre story. Both are reminders that inference cost is now a physical constraint measured in milliseconds, joules and euros, not only in dollars per million tokens.
The week's through-line is that no single number settles it. Magnitude's 92% decode claim is Metal-specific and self-reported. Quail's billion tokens per minute comes from one multi-join query on one H100, with the team itself flagging two benchmark queries as areas for improvement. EnerTune's savings are modelled against baselines in a paper. The pricing index is a frozen snapshot of list prices that most buyers do not pay.
What is consistent is the direction of the engineering: tune kernels to the exact chip, plan queries around the model, and treat power as a first-class metric. Whether that shows up as cheaper inference depends less on any single benchmark than on how much of it reaches production.
Sources
9- 01Launch HN: Magnitude (YC S25) - Self-optimizing inference engine for agentsEN
- 02Jointly optimizing SQL queries and LLM inference for up to 14x speedupsEN
- 03Hitting 1B tokens/minute on 1 GPU combining a query planner and inference engineEN
- 04AI API pricing index - September 2026EN
- 05Beyond Utilization: Energy-Conscious GPU Sharing for Inference ServingEN
- 06Singapore, New Zealand operators seek partnerships to cut costsEN
- 07The pricing calculator made people more confused about pricingEN
- 08Mind-1: Cutting robot inference latency from 82ms to 32msEN
- 09Driving An EV Now Costs Half As Much As Diesel, New Analysis ShowsEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.