OpenAI's $500 ChatGPT Tier Lands Days After $1.3 Median Input Price Snapshot
OpenAI announced a $500-a-month ChatGPT subscription at its DevDay on Tuesday, the same week a month-end snapshot put the median list price for an input million tokens at $1.32 across 81 public records.

OpenAI used its Tuesday DevDay to announce a new top subscription tier, Pro 500, at $500 a month, according to Business Insider. That is $300 above the company's previous top plan. For a single seat, it works out to $6,000 a year.
The tier bundles access to what OpenAI calls "Ultrafast" computing for GPT-6 Astra in ChatGPT Work and Codex. Business Insider reports the option promises an 8x speed increase inside Codex and carries the highest included usage the company offers. The same report says OpenAI is reopening its older $200-a-month Pro plan, which it paused on September 10, but cutting that plan's compute allowance in half. The allowance drops to 10x the $20 Plus tier from a previous 20x multiple, and weekly chat messages in GPT-6 Pro fall from 200 to 100.
Thibault Sottiaux, the engineering lead for OpenAI's Codex, wrote on X on Monday that the changes net out at half the dollar in API spend compared with the prior plan. Business Insider quotes him directly. No other outlet in this dossier confirms the pricing details, and OpenAI has not published a public price list that this piece could verify independently.
The timing matters because a September snapshot of published API pricing shows the list-price floor is nowhere near $500 a month for most workloads. AICostBudget's AI API Pricing Index, frozen at a 2026-09-30 UTC cutoff, covers 81 public pricing records across 11 providers. Across 69 records with a current input-token scalar, the observed range runs from $0.10 to $30 per million tokens. Across the same 69 records for output tokens, the range runs from $0.10 to $180 per million.
Discounts, caching and the sticker price
The index's headline medians are $1.32 per million input tokens and $6 per million output tokens, giving a median output-to-input ratio of 5.0x. Those numbers are list prices, and the index is explicit that list prices are not what most buyers pay. Across 58 comparable records, the observed cached-input discount ranges from 75% to 98%, with a median of 90%. Across 44 Batch-capable records, the observed batch input discount ranges from 20% to 50%, with a median of 50%.
That gap is the whole argument for the inference-engine and query-optimization projects that appeared in the last week. If a workload can move its repeated prefixes into cache and its bulk jobs into batch, the effective price per token falls well below the median. If it cannot, the list price is the bill.
Two projects published on 30 September attack the problem from opposite ends of the stack. On the database side, the Full Stack Data Lab's QUAIL work, posted on 30 September, jointly optimizes SQL query planning and LLM inference. Its BIO-4 benchmark query runs over 5,000 long medical reports and 4,144 reaction terms, with the reaction list used twice across two joins. The authors argue that sending millions of related model calls to a general-purpose engine such as vLLM as separate requests carries a large cost, because each prompt is rendered and served as an independent inference request.
Modal's write-up of the same system, also dated 30 September, puts numbers on it. Quail processes over a billion tokens per minute per H100 GPU on one multi-join query, which Modal says is more than 10x faster than its vLLM baseline on identical hardware, and costs under 6 cents per billion tokens on Modal. Across the newly released AI-SQL benchmark, Modal reports Quail at 1.84x faster than vLLM, geometrically averaged over tasks, including two queries the authors designed to show where AI-SQL inference still needs work.
The 1.84x geometric mean and the 10x single-query figure come from the same team and should not be read as the same measurement.
Local kernels and robot latency
On the client side, Magnitude, a Y Combinator S25 company, launched an open-source inference engine on 30 September that compiles and tunes its kernels on the user's own device before a model runs. Its GitHub README claims open models run up to 2x faster than llama.cpp, with 92% faster decode on Metal and 19% on CUDA, and 27% less memory per agent. It supports Apple Silicon, NVIDIA, AMD, or CPU only. Those figures are the project's own and have not been independently reproduced in this dossier.
Robot inference is chasing the same clock. MindOn introduced Mind-1 on 30 September, claiming it cuts robot inference latency from 82ms to 32ms. The company notes that at an end-effector speed of 3 m/s, 10 ms of latency corresponds to roughly 3 cm of travel, enough to affect precise grasping, insertion, or contact-rich manipulation. Its speed claims cover logistics and everyday human environments and come from MindOn's own testing.
Academic work points at a cost that utilization metrics hide. A paper presented at SOSP '26 on 28 September by Prasoon Sinha, Dimitrios Liakopoulos, Nathan Lemma and Neeraja J. Yadwadkar describes EnerTune, a serving system that models per-model power and the power draw of colocated models on shared GPUs. The authors report energy reductions of 1.4x to 2.3x and power-draw reductions of 1.3x to 2.6x over state-of-the-art baselines while meeting performance SLOs. Their argument is that optimizing purely for utilization can raise energy consumption, and that profiling every configuration at scale costs too much energy to be practical.
Where the money actually lands is still contested. Light Reading reported on 30 September that StarHub and M1 are in merger talks in Singapore, and that 2degrees and OneNZ in New Zealand have proposed combining their radio networks. StarHub cautions negotiations are continuing. Light Reading notes that a combined StarHub and M1 would be roughly Singtel's scale, and that Singtel reportedly held a 43% mobile share in June, with M1 at 22%, StarHub at 21% and Simba at 14%.
One pricing lesson from the same week has nothing to do with GPUs. Pinecone removed its pricing calculator after finding that a single misinterpreted input could overstate an estimate by as much as 1,000x, according to a first-person account by Gkogan published on 30 September. Visitors who did not see the calculator were 16% more likely to sign up and 90% more likely to contact the company, with no rise in pricing support tickets. Seven in ten employees had expected the calculator version to win.
Sources
9- 01ChatGPT's new plan costs $500 a monthEN
- 02AI API Pricing Index - September 2026EN
- 03Hitting 1B tokens/minute on 1 GPU combining a query planner and inference engineEN
- 04Jointly optimizing SQL queries and LLM inference for up to 14x speedupsEN
- 05Launch HN: Magnitude (YC S25) - Self-optimizing inference engine for agentsEN
- 06Mind-1: Cutting robot inference latency from 82ms to 32msEN
- 07Beyond Utilization: Energy-Conscious GPU Sharing for Inference ServingEN
- 08Singapore, New Zealand operators seek partnerships to cut costsEN
- 09The pricing calculator made people more confused about pricingEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.