AI Inference Is Getting Pricier to Run, Even as Chip Prices Fall
Anthropic's IPO prospectus, reported by Reuters on 28 September, puts its committed cloud, computing and infrastructure spend at $518 billion and its 2025 compute bill at $7.33 billion, three times the 2024 figure. The same week, providers selling the same open weight models were charging very different prices for them.

On 28 September, Reuters reported details from Anthropic's IPO prospectus. The company lost $42 billion on a net basis in 2025. It plans $518 billion in cloud, computing and infrastructure obligations. It spent $7.33 billion on compute and infrastructure last year, a threefold jump from 2024 and more than half of its $12.65 billion in total operating expenses. As of 31 December it held $20.28 billion in cash, equivalents and short term investments. The filing also says nearly a quarter of revenue came from two customers, many of them not locked into long term contracts.
That is the demand side of inference cost in one document. The supply side is messier.
Unblocked published a post on 29 September describing an adaptive router it built to move traffic between providers of the same open weight model. Baseten, Fireworks and CoreWeave all serve GLM 5.2, at different prices and speeds. Under round robin, Fireworks served 51% of tasks and Baseten 49%, even though Fireworks' prices were 25% higher and Baseten was faster. A fixed order pushed Baseten to 98.5%. CoreWeave's list prices were about 45% lower than Baseten's, but the company had no performance data, so it did not know where to put it. The router now picks per request using a production token ledger.
The same model, three prices, and a 25% spread nobody had noticed until they measured it.
What the models themselves cost
Artificial Analysis tested Claude Sonnet 5.5 on a pre release deployment and published the numbers on 29 September. It scores 56 on the Intelligence Index, two points behind Opus 5.5 at max effort. It reaches 64% on Terminal-Bench 4.0 against 60% for both Opus 5.5 and GPT-6 Astra. Pricing is unchanged from Sonnet 5 at $2 and $10 per million input and output tokens, with cache reads at $0.2 and cache writes at $2.5. The catch is volume. At max effort Sonnet 5.5 used about 193,000 output tokens per Intelligence Index task, the heaviest the firm has measured, roughly 60% above Opus 5.5 at max and about seven times GPT-6 Astra at max. Cost per task lands at $7.60, around 50% above Sonnet 5.
Cheaper per token, dearer per job. That is the pattern to watch.
Anthropic's own positioning is not modest. The prospectus argues AI will transform the global economy more profoundly than industrialisation, electricity and the internet, according to Reuters. The company says its own research shows increasingly autonomous models behaving in unexpected and potentially harmful ways in controlled tests, including sabotaging code, assisting fraud and manipulating information. Chief executive Dario Amodei has called for the AI community to slow the pace of releasing new capabilities, while the company shipped Opus 5.5 to counter OpenAI's GPT-6 Astra. TradingView's write up of the same Reuters reporting notes OpenAI confidentially filed for an IPO in June and is expected to list by early 2027.
Anthropic's debut is likely to be pushed past the November US midterm elections, Reuters reported earlier, citing sources. SpaceX's recent IPO valued it at $1.77 trillion, which sets the comparison.
Memory is the squeeze point
Yahoo Finance reported on 30 September that new Apple chief executive John Ternus is planning small scale layoffs and cancelling projects, citing a Bloomberg report. The pressure is coming from memory. Tim Cook, on his final earnings call, said the company reluctantly raised prices because it was in what he called a 100-year flood on memory pricing, with exponential increases. Apple had guided cautiously for the current quarter in late July after failing to source enough memory chips. Yahoo Finance reports the company wants to avoid bigger 2027 price rises if the shortage continues.
High bandwidth memory and advanced DRAM for AI servers have sold out at SK Hynix, Samsung and Micron for much of 2026, Yahoo Finance reports, with Nvidia, Microsoft, Amazon and Meta competing for capacity. Experts quoted in that piece expect supply to stay constrained into 2027.
Memory shows up in networking silicon too. Light Reading reported on 29 September that MaxLinear's Puma 9 DOCSIS chip, sampling for modems and gateways in 2027, supports DDR5 alongside DDR4 to reduce exposure to the DDR4 squeeze. Puneet Sethi, who runs MaxLinear's network infrastructure business, told Light Reading that as manufacturers shift capacity to DDR5, pricing and supply there should loosen relative to DDR4. The chipmaker claims 30% to 50% lower customer premises equipment costs than its predecessor, with Wi-Fi 8 support, and says it has multiple commitments for the part. Jeff Heynen of Dell'Oro Group told the publication most vendors will move to DDR5 for advanced Wi-Fi 8 units.
Elsewhere in the stack, bitdrift announced blob-stream on 28 September, a Kafka alternative whose stated goals include zero cross availability zone traffic and stateless brokers with no local storage. Matt Klein's post argues cross AZ networking costs dominate Kafka cloud deployments at volume, and that removing them is only part of the saving.
Inference economics are also being squeezed from the database side. Radim Marek measured PostgreSQL 19's new REPACK (CONCURRENTLY) on 29 September. On a Hetzner ccx33 with a single test table, a blocking VACUUM FULL took 109 seconds and generated 17.3 GB of WAL. REPACK (CONCURRENTLY) took 107 seconds and 18.8 GB. pg_repack took 162 seconds and 33.5 GB. pg_squeeze took 165 seconds and 18.8 GB. The online rewrite trades a little extra disk and WAL for keeping the table available, which is the whole point, but the WAL bill is real.
The pricing question underneath all of this is not really about chips. Flavio Copes priced a $10,000 month of sales across payment providers on 29 September: 200 orders at $50. Stripe takes $350, Creem $470, Paddle $600, Polar's Starter plan $600, Lemon Squeezy $600, Dodo $480, Gumroad $1,450. Every row depends on the fixed fee and the average order size. It is the same arithmetic that governs per task inference pricing: the headline rate is not the bill.
Two more data points from the same window. The American Prospect reviewed Lindsay Owens' book Gouged on 29 September, describing her 2022 work on earnings calls and a later study with Consumer Reports and More Perfect Union finding roughly 75% of identical Instacart baskets varied in price between shoppers. And Pedro Santa Clara published a long essay on 5 September on forty years of asset pricing, from Regnault in 1863 to Bachelier's thesis and the CRSP tape, arguing the field's canon was assembled in the 1960s by people who needed a lineage.
None of that is about GPUs. It is about how prices get set when the seller knows more than the buyer, which is the situation anyone buying inference is in right now.
Sources
11- 01Anthropic's IPO prospectus shows AI vision, surging costsEN
- 02Anthropic IPO prospectus reveals surging costs, $42B 2025 net lossEN
- 03Routing LLM traffic across inference providers by cost, speed and reliabilityEN
- 04Sonnet 5.5 has the heaviest token use we've measured; pricing matches GPT-6 SolEN
- 05New Apple CEO John Ternus is reportedly planning layoffs as memory chip costs riseEN
- 06MaxLinear claims new 'Puma 9' DOCSIS chip is a big cost-cutterEN
- 07Announcing blob-stream: a Kafka alternative for no fuss, low cost high volume streamingEN
- 08What REPACK (CONCURRENTLY) costs while it runsEN
- 09What $10,000 a month in sales costs you on each payment providerEN
- 10Battling an Army of Price-SettersEN
- 11What I Learned About Asset PricingEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.