Inference cost moves to the centre of the AI stack as new engines ship
Magnitude, a Y Combinator S25 startup, published its open source inference engine on GitHub on 30 September, claiming open models run up to 2x faster than llama.cpp on the same hardware.

Magnitude published its repository on 30 September at 17:37 UTC, according to the GitHub listing. The pitch is narrow: an inference engine for agents that compiles and tunes kernels on the device it runs on. The company says that produces up to 2x faster performance than llama.cpp, with 92% faster decode on Apple's Metal backend and 19% on CUDA.
The claim is a vendor benchmark. Nobody outside Magnitude has reproduced it yet. The repository does not name the models, prompt lengths or batch sizes behind the numbers. Treat the 2x as a marketing figure until someone independent runs it.
Where the company aims matters more. Magnitude says it ships a desktop app that connects to agents people already use, listing Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi and Cline. Anything else goes through an OpenAI-compatible API. The engine runs on Apple Silicon, NVIDIA or AMD GPUs, or a CPU alone, per the repo. Magnitude also claims 27% less memory per agent, freed when agents stop, and says sessions share prefix caches to avoid slowdowns under concurrency.
The same week, a different layer of the stack
Two hours later on 30 September, a second release landed. Full Stack Data Lab published a blog post on Quail, a query-aware inference layer built with Modal, and Modal put up its own write-up at 17:38 UTC the same day. Quail attacks AI-SQL, the practice of calling language models row by row from inside SQL queries.
The numbers are large. Full Stack Data Lab says a single query can generate hundreds of thousands or millions of model calls, because an AI function evaluates one prompt per row, and a naive join evaluates one prompt per pair of rows. On a benchmark query called BIO-4, which joins 5,000 medical reports against 4,144 reaction terms, the lab walks through the cost of sending those calls to a general-purpose engine such as vLLM as separate requests.
Modal's post reports the result: Quail processes over a billion tokens per minute per H100 GPU on one multi-join query, more than 10x faster than its vLLM baseline on the same hardware, at under 6 cents per billion tokens on Modal. Across a newly released benchmark, Modal says Quail runs 1.84x faster than vLLM, geometrically averaged over tasks. It flags that the set includes two queries designed to show where AI-SQL inference still needs work.
I see it as a point on the LLM pareto optimal curve in a regime that had a large revealed latent demand (no thinking, single token, low latency acceptable intelligence) that was under-invested into because of a race to higher intelligence. Karpathy-san, on Jev
That quote appears in Modal's post, attributed there to Karpathy, commenting on TypeSafe AI's Jev model. The argument is that a lot of inference demand sits at the cheap, low-latency end and has been under-served because labs chased frontier capability instead. Both Magnitude and Quail are bets on that gap.
Neither write-up is peer reviewed. Both are company blogs accompanying code releases, and the comparison baselines, one against llama.cpp, one against vLLM, are chosen by the authors. The direction of travel is consistent, but the magnitudes should not be treated as settled.
What the electricity and the price list say
Academic work published this week points at a related problem that faster tokens do not solve on their own. A paper presented at SOSP '26, dated 28 September and listed in the ACM proceedings, argues that optimising GPU clusters purely for utilisation can increase energy consumption. The authors, Prasoon Sinha, Dimitrios Liakopoulos, Nathan Lemma and Neeraja J. Yadwadkar, describe a system called EnerTune that models per-model performance and power analytically instead of profiling every configuration, and packs models onto shared GPUs accordingly. They report reducing energy consumption by 1.4 to 2.3 times and power draw by 1.3 to 2.6 times over state-of-the-art baselines while still meeting performance targets.
Then there is the list price side. AICostBudget published its September pricing index on 30 September, a frozen snapshot taken at a 2026-09-30 UTC cutoff, covering 81 public pricing records across 11 providers. The median input price across 69 records is $1.32 per million tokens and the median output price is $6 per million, a ratio of 5.0x. Output pricing, not prompt ingestion, is usually the larger component.
That index also undercuts a common shortcut in cost discussions. Of 58 comparable cached-input records, the median discount is 90%. Of 44 batch-priced records, the median discount is 50%. Any workload comparison that ignores caching and batch pricing is describing a bill most customers will never see, and the index says so directly. On the provider side, its medians run from Mistral AI at $0.20 input and $0.60 output per million tokens across 7 records, and DeepSeek at $0.30 and $1.20 across 3, up to Anthropic at $5 and $25 across 13 records and OpenAI at $2 and $11 across 18. These are medians of public list prices, not effective prices after discounts, and the populations differ in size, so they are not a like-for-like model ranking.
Speed is only one of the costs
Outside the data centre, latency has a measurable physical cost. MindOn introduced its Mind-1 physical AI model on 30 September, saying it cuts robot inference latency from 82ms to 32ms and pushes manipulation tasks to human-level cycle times. The company's own explanation of why that matters is the useful part: at an end-effector speed of 3 metres per second, 10ms of latency corresponds to roughly 3cm of movement, enough to affect precise grasping or insertion. MindOn also points at a data problem, noting that much robot manipulation data comes from teleoperation that is slower than natural human motion, so models learn a slower tempo along with the task.
There is a consumer-facing cost story in the same window. CleanTechnica reported on 29 September that Transport & Environment analysis puts EV running costs at less than half of diesel per kilometre in Europe, based on average consumption of 6.9 litres per 100km for diesel and 20.2 kWh per 100km for battery electric, with electricity at €0.343 per kWh. T&E's Antony Froggatt said in the post that diesel drivers are paying €32 more per tank than at the start of the year and warned of supply risk if the US blocks diesel exports. T&E notes its charging estimate is likely an upper bound.
The pricing page itself can be a cost. A first-person account published on 30 September by Gokhan Kogan, describing his time at Pinecone, reports that a usage calculator on the pricing page was producing estimates overstated by as much as 1,000x after a single misinterpreted input. In an A/B test, he writes, visitors who did not see the calculator were 16% more likely to sign up and 90% more likely to contact the company, with no increase in pricing support tickets. He also notes that 7 in 10 employees polled internally expected the calculator version to perform better. That is one company's experience, not a general law, but it is a reminder that perceived cost and actual cost diverge easily.
Capacity economics are changing at the network layer too. Light Reading reported on 30 September that StarHub and M1 are in merger talks in Singapore, where they already share 5G spectrum and radio access network through a jointly owned company called Antina. Singtel held 43% of the mobile market in June, with M1 at 22%, StarHub at 21% and Simba at 14%, according to the report. In New Zealand, 2degrees and OneNZ have proposed combining their radio networks, with a decision from the Commerce Commission expected next year.
The through-line is not that inference is about to get free. The optimisations arriving this week are aimed at specific bottlenecks: agent decode speed on consumer hardware, row-by-row SQL calls, and idle GPU power. Each comes with its own benchmark and its own vendor. The pricing index suggests the list prices themselves have not moved much. The savings are being sold at the engine layer, and buyers will have to test them on their own workloads.
Sources
9- 01Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agentsEN
- 02Jointly optimizing SQL queries and LLM inference for up to 14x speedupsEN
- 03Hitting 1B tokens/minute on 1 GPU combining a query planner and inference engineEN
- 04AI API pricing index - September 2026EN
- 05Mind-1: Cutting robot inference latency from 82ms to 32msEN
- 06The pricing calculator made people more confused about pricingEN
- 07Driving An EV Now Costs Half As Much As Diesel, New Analysis ShowsEN
- 08Beyond Utilization: Energy-Conscious GPU Sharing for Inference ServingEN
- 09Singapore, New Zealand operators seek partnerships to cut costsEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.