Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Inference costs split in two: cheaper per token, pricier per GPU

Magnitude, a Y Combinator S25 startup, published an open source inference engine on 30 September that it says runs open models up to 2x faster than llama.cpp on Apple Silicon, NVIDIA, AMD or a plain CPU, with no token costs. The same day, a price index covering 81 public records put the median listed output price at $6 per 1M tokens against $1.32 for input.

AI & modelsExplainerGrace OkonkwoPublished: 30 September 20266 min readSources 10
Inference costs split in two: cheaper per token, pricier per GPU

Magnitude's repository went up on 30 September. The pitch is straightforward: the engine compiles and tunes its kernels on your device before a model runs, rather than shipping precompiled binaries for broad hardware classes. The company reports 92% faster decode on Metal and 19% on CUDA against llama.cpp, plus 27% less memory per agent, freed when agents stop. The licence is Apache 2.0.

That is one answer to the cost problem: stop paying per token.

It is not the only one, and it is not the one the rest of this week's material points at. Quail, an inference engine built by the Full Stack Data Lab team behind DocETL, targets the other half of the bill. Its authors argue that AI-SQL queries, where an LLM call runs per row or per row pair, are badly served by general-purpose engines. Sending millions of related model calls to vLLM as separate requests carries a large cost, they write. Quail instead plans the query and the inference together.

The numbers are specific. On a multi-join query where planning matters most, Quail hits over a billion tokens processed per minute on one H100 GPU, which the team says is more than 10x its vLLM baseline on the same hardware. On Modal that works out to under 6 cents per billion tokens. On the team's own AI-SQL benchmark the gain is smaller: 1.84x faster than vLLM, geometrically averaged, including two queries they designed to show where AI-SQL inference still falls short.

Per-token prices are not the whole bill

AICostBudget's September index, frozen at the 30 September UTC cutoff, covers 81 public pricing records across 11 providers. Across 69 records with a current input-token scalar the range runs from $0.10 to $30 per 1M tokens. Output runs from $0.10 to $180. The median input price is $1.32 and the median output price is $6, a ratio of 5.0x on the 69 records that carry both.

Provider medians diverge sharply. OpenAI's 18 input records and 18 output records median at $2 and $11. Google Gemini, on 14 each, medians at $0.75 and $4.125. Anthropic's 13 records median at $5 and $25. xAI's nine sit at $1.25 and $2.50, Mistral's seven at $0.20 and $0.60, DeepSeek's three at $0.30 and $1.20.

The index also flags what a naive comparison misses. Across 58 comparable records the median cached-input discount is 90%, and across 44 Batch-capable records the median Batch input discount is 50%. The report's own framing is blunt: a workload comparison that ignores cache and Batch pricing omits published savings modes for a large share of the tracked records. It adds that document pages, storage, session duration and other time-based units cannot be safely flattened into a token-only table, which is a fair warning for anyone building a single cost-per-token league table.

The pricing data is a snapshot, not a market clearing price. GPU capacity is moving separately.

Where the GPU money is going

Business Insider reported on 30 September that OpenAI announced a $500-a-month ChatGPT tier at its Tuesday DevDay, a $300 increase over the company's previous most expensive plan. The tier includes "Ultrafast" computing for GPT-6 Astra in ChatGPT Work and Codex, which OpenAI says is an 8x speed increase in Codex, plus the company's highest included usage. Business Insider also reported that OpenAI is reopening its $200 Pro plan while halving the compute allowance, from 20x the $20 Plus plan to 10x, and cutting GPT-6 Pro chat messages from 200 to 100 a week. Thibault Sottiaux, the engineering lead for Codex, said in a Monday X post that the changes net out at half the dollar in API spend compared to the prior plan.

On the hardware side, the SOSP '26 paper from Prasoon Sinha, Dimitrios Liakopoulos, Nathan Lemma and Neeraja J. Yadwadkar, published 28 September, makes an argument that cuts against the usual utilisation metric. GPUs are expensive, yet inference clusters remain heavily underutilised, so systems adopt multiplexing. But the authors write that optimising solely for utilisation can counterintuitively increase energy consumption. Their system, EnerTune, uses analytical models of per-model performance and power, including the draw of colocated models on shared GPUs, inside an energy-aware bin-packing algorithm. It reports 1.4x to 2.3x lower energy consumption and 1.3x to 2.6x lower power draw than state-of-the-art baselines while meeting performance SLOs.

That is a cost claim about a bill most token-price tables do not show.

Robotics is running the same play in a different setting. MindOn's Mind-1, announced 30 September, cuts inference latency from 82ms to 32ms, according to the company's own write-up. The post explains why that matters: at an end-effector speed of 3 m/s, 10ms of latency corresponds to roughly 3cm of movement, enough to affect precise grasping, insertion or contact-rich manipulation. MindOn also points at training data tempo, arguing that teleoperated demonstrations are slower than natural human motion and that models learn the tempo as well as the task.

Two bills, two strategies

Pinecone's experience with a pricing calculator, recounted by Gabriel Kogan on 30 September, is a reminder that the cost number customers see is itself a product decision. The calculator produced estimates overstated by as much as 1,000x after one misinterpretation, and gave users false confidence that discouraged them from checking the docs. When the company A/B tested removing it, visitors who did not see the calculator were 16% more likely to sign up and 90% more likely to contact the company, with no rise in pricing support tickets. In an internal poll, 7 of every 10 employees expected the version with the calculator to do better.

Two more data points sit outside the AI stack but frame the same argument about substituting capital for fuel. Transport & Environment, analysed by CleanTechnica on 29 September, says EV drivers pay less than half what diesel drivers pay per kilometre, based on 6.9 L/100 km for diesel, 20.2 kWh/100 km for battery electric and an electricity price of €0.343/kWh. The group's Antony Froggatt said diesel drivers are paying €32 more per tank than at the start of the year, with around €16 of that extra refinery margin.

Light Reading reported on 30 September that StarHub and M1 are in merger talks in Singapore, with the combined entity roughly matching Singtel's scale, and that New Zealand's 2degrees and OneNZ have proposed combining their radio networks. The China Telecom and China Unicom 5G arrangement, the largest such sharing deal, covers 1.5 million basestations and claimed $56 billion in capex savings. None of that is AI inference. All of it is the same trade: share the fixed asset, cut the marginal cost.

For buyers, the practical split is now clear. Per-token prices are falling in the published tables, and caching and Batch modes cut them further. Per-GPU-hour costs are a different market, and the research this week suggests most of it is still wasted. Neither Magnitude's local kernels nor Quail's query planner fixes that. They just move where the meter sits.

Comments 0

Sources

10
  1. 01Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agentsEN
  2. 02Hitting 1B tokens/minute on 1 GPU combining a query planner and inference engineEN
  3. 03Jointly optimizing SQL queries and LLM inference for up to 14x speedupsEN
  4. 04AI API pricing index - September 2026EN
  5. 05ChatGPT's new plan costs $500 a monthEN
  6. 06Beyond Utilization: Energy-Conscious GPU Sharing for Inference ServingEN
  7. 07Mind-1: Cutting robot inference latency from 82ms to 32msEN
  8. 08The pricing calculator made people more confused about pricingEN
  9. 09Driving An EV Now Costs Half As Much As Diesel, New Analysis ShowsEN
  10. 10Singapore, New Zealand operators seek partnerships to cut costsEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.