Inference Costs Split: $500 ChatGPT Tier, 2x Faster Open Engines, 6c per Billion Tokens
OpenAI opened a $500-a-month ChatGPT Pro tier on 29 September, its most expensive plan yet and a $300 jump over the previous top tier, according to Business Insider, as cheaper open inference stacks race to undercut it.

The new plan, announced at the company's Tuesday DevDay, bundles access to "Ultrafast" computing for GPT-6 Astra in ChatGPT Work and Codex. Business Insider reported that an 8x speed increase in Codex is promised for that mode. The same report says OpenAI is reopening its $200-a-month Pro plan, which it paused on 10 September, but halving the compute allowance to 10x the $20 Plus tier instead of 20x, and cutting GPT-6 Pro chat messages from 200 to 100 a week.
Thibault Sottiaux, the engineering lead for Codex, wrote in a Monday post on X that the changes net out at half the dollar in API spend compared with the prior plan, according to Business Insider.
"I wanted to make sure to share this change ahead of time so you can all understand it before we shower you with good news," he wrote.
That is the top of the market. The bottom is moving faster.
Open engines and self-tuning kernels
On 30 September, YC-backed Magnitude published an open source inference engine that compiles and tunes kernels on the user's own hardware before a model runs. The company claims open models run up to 2x faster than llama.cpp. The GitHub README states 92% faster decode on Apple Silicon's Metal backend and 19% on CUDA, plus 27% less memory per agent and shared prefix caches for concurrent sessions.
It ships as a desktop app under Apache 2.0, supports Apple Silicon, NVIDIA, AMD or CPU-only, and connects to Pi, OpenCode, Hermes, Codex and other agents. Two hours earlier, Modal and the Full Stack Data Lab published a joint SQL-plus-inference engine called Quail. Modal's write-up claims Quail hits over a billion tokens processed per minute on one H100 GPU on a multi-join query, more than 10x its vLLM baseline on the same hardware, at under 6 cents per billion tokens on Modal. Across the team's new AI-SQL benchmark, the geometric mean speedup over vLLM is 1.84x, and the post names two queries designed to show where the approach still falls short.
The authors are explicit that generic engines are the wrong shape for this workload. In the Full Stack Data Lab post, Shreya Shankar, Charles Frye, Fergus Finn, Arnav Dhariya, Joseph Barrow and Meryem Arik write that sending millions of related model calls to an engine like vLLM as separate requests carries a large cost. Their example query, BIO-4 in the QUAIL-B benchmark, filters 5,000 long medical reports against 4,144 reaction terms used twice, so a naive execution turns one SQL statement into a very large number of model calls.
The price floor is published, and it is not flat
AICostBudget's September pricing index, frozen at a 2026-09-30 UTC cutoff, covers 81 public records across 11 providers. It puts the median input price at $1.32 per million tokens and the median output price at $6, with output priced 5.0x input at the median across the 69 records that carry both numbers. The observed ranges run from $0.10 to $30 per million input tokens and $0.10 to $180 per million output tokens.
The same dataset gives provider-level medians: OpenAI at $2 input and $11 output across 18 records, Google Gemini at $0.75 and $4.125 across 14, Anthropic at $5 and $25 across 13, xAI at $1.25 and $2.5 across nine, Mistral AI at $0.20 and $0.60 across seven, DeepSeek at $0.30 and $1.2 across three, and Moonshot AI at $0.95 and $4 across three. Across 58 comparable cached-input records the median discount is 90%, and across 44 batch-priced records the median batch discount is 50%. A comparison that ignores either mode is comparing list prices, not bills.
Hardware is where the argument gets harder to settle. EnerTune, a paper by Prasoon Sinha, Dimitrios Liakopoulos, Nathan Lemma and Neeraja J. Yadwadkar at SOSP '26 on 28 September, argues that multiplexing GPUs purely to raise utilization can raise energy use. The authors report their energy-aware bin-packing system meets performance SLOs while cutting energy consumption by 1.4 to 2.3x and power draw by 1.3 to 2.6x against state-of-the-art baselines.
Not all of this week's cost news is about data centres. Transport & Environment data analysed on 29 September puts EV running costs at less than half of diesel per kilometre, based on 6.9 litres per 100 km for diesel and 20.2 kWh per 100 km for battery electric at an electricity price of EUR 0.343 per kWh. Diesel drivers are paying EUR 32 more per 50-litre tank than at the start of the year, about EUR 16 of it extra refinery margin, per ECB estimates cited by CleanTechnica.
Two lessons from the week sit outside AI entirely. Pinecone removed the pricing calculator from its pricing page after an A/B test found visitors who did not see it were 16% more likely to sign up and 90% more likely to contact the company, with no rise in pricing support tickets, according to founder Greg Kogan. And in telecoms, Light Reading reported on 30 September that Singapore's StarHub and M1 are in merger talks while New Zealand's 2degrees and OneNZ have proposed combining their radio access networks.
The through-line: when a unit of work gets cheap enough, buyers stop reading the calculator and start measuring the bill.
Sources
9- 01ChatGPT's new plan costs $500 a monthEN
- 02Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agentsEN
- 03Hitting 1B tokens/minute on 1 GPU combining a query planner and inference engineEN
- 04Jointly optimizing SQL queries and LLM inference for up to 14x speedupsEN
- 05AI API pricing index - September 2026EN
- 06Beyond Utilization: Energy-Conscious GPU Sharing for Inference ServingEN
- 07Driving An EV Now Costs Half As Much As Diesel, New Analysis ShowsEN
- 08The pricing calculator made people more confused about pricingEN
- 09Singapore, New Zealand operators seek partnerships to cut costsEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.