Inference economics shift as Anthropic's $7.33B compute bill meets Sonnet 5.5 token surge
Anthropic's IPO prospectus, seen by Reuters, shows the company spent $7.33 billion on compute and infrastructure in 2025, a threefold rise from 2024, while one of its newest models turns out to burn more output tokens per task than any model Artificial Analysis has measured.

On 29 September, Artificial Analysis published an evaluation of Anthropic's Claude Sonnet 5.5. The firm found it uses roughly 193,000 output tokens per Intelligence Index task at max effort. That is the highest token consumption the testing firm has recorded, around seven times that of GPT-6 Astra at max effort. It is a measurement, not a funding round or a chip launch, and it lands at the same time as the industry's biggest cost disclosure of the year.
Anthropic's IPO prospectus, first reported by Reuters and covered by CNBC and TradingView on 28 and 29 September, shows the company spent $7.33 billion on compute and infrastructure last year, more than half of its $12.65 billion total operating expenses. It also plans $518 billion in cloud, computing and infrastructure obligations in coming years. Per-token prices have become the standard way to compare models. Artificial Analysis notes that Anthropic priced Sonnet 5.5 identically to Sonnet 5 at $2 per million input tokens and $10 per million output tokens, matching GPT-6 Sol. The same evaluation says the model costs about $7.60 per task, roughly 50% more than Sonnet 5's cost per task, because it simply emits more tokens to reach comparable scores.
Token use is the cost variable nobody prices at the counter
At max effort, where it reaches performance nearing that of Opus 5.5, Claude Sonnet 5.5 used ~193k Output Tokens per Intelligence Index Task. This is the highest token use we have measured.
Artificial Analysis also flags a caveat: the evaluations ran on a pre-release deployment with a bug affecting structured outputs, fixed for the public release, and the firm expects to re-run relevant tests. Sonnet 5.5 reaches 56 on the Intelligence Index, two points behind Opus 5.5 at max effort, and scores 64% on Terminal-Bench 4.0 against 60% for Opus 5.5 and GPT-6 Astra.
Meanwhile the bill for serving that traffic keeps moving in one direction. According to a post on getunblocked.com dated 29 September, the company routes open-weight GLM 5.2 traffic across Baseten, Fireworks and CoreWeave, and found CoreWeave's list prices about 45% below Baseten's. Its adaptive router weighs cost at 0.7 and speed at 0.3, on the logic that the model is identical whichever provider serves it. The post is candid about the earlier setup: a round-robin split sent 51% of tasks to Fireworks at prices 25% higher than Baseten's, for no benefit.
Memory is where the money is going
On the hardware side, the pressure has not eased. Yahoo Finance reported on 30 September that new Apple CEO John Ternus is planning small-scale layoffs within large teams and cancelling projects, citing a Bloomberg report, as memory costs bite. The article quotes Apple's outgoing CEO Tim Cook on his final earnings call describing what he called a 100-year flood on memory pricing, with exponential increases. Apple had already issued cautious revenue guidance in late July because it could not source enough memory chips.
The same piece notes that SK Hynix, Samsung Electronics and Micron have largely sold out premium AI memory capacity through much of 2026, as Nvidia, Microsoft, Amazon and Meta build out AI infrastructure. Supply is expected to stay constrained into 2027. That is the backdrop against which every inference cost projection now has to be read: it is not only about how many tokens a model emits, but what the memory inside the serving cluster costs.
Anthropic's own filing, as reported by Reuters, puts the demand side in stark terms. Revenue grew twelvefold in 2025 to nearly $4.6 billion, but the company posted a net loss of $42 billion, including a roughly $34 billion accounting charge tied to the rising estimated value of financing that could convert into shares. It also warned that nearly a quarter of revenue came from two customers, and that many of its largest clients were not locked into long-term contracts.
A pricing question with no clean answer
The gap between per-token sticker prices and real cost per task is not a small accounting detail. Artificial Analysis places Sonnet 5.5 off the Intelligence versus Cost per Task Pareto frontier, with lower effort settings sitting behind GPT-6 Sol configurations that deliver equivalent performance for fewer output tokens. The high effort setting is described as very narrowly behind GPT-6 Sol on intelligence at effectively the same cost per task.
Other sources in the dossier point in the same direction, though they are not about inference directly. A 29 September post on flaviocopes.com compares payment provider fees on a $10,000 month and finds effective rates clustered between 3.5% and 6% for everyone except Gumroad at 14.5%. It is a reminder of how much apparent price comparison depends on the unit you choose, a $0.40 fixed fee is 0.8% of a $50 order and 4% of a $10 order.
Whether inference pricing follows the same pattern is the open question. Anthropic's prospectus implies a company betting that scale will eventually bend the cost curve, while its own model releases are making tasks more expensive to serve. The IPO is likely to be pushed to after the November US midterm elections, Reuters previously reported, citing sources, with an expected valuation above $2 trillion, more than double its own $965 billion estimate in May.
Investors assessing that number will have to weigh two documents published a day apart: a prospectus showing $518 billion in future infrastructure commitments, and a benchmark showing the company's newest mid-tier model uses more output tokens than anything previously measured. Neither is fatal on its own. Together they frame the question the AI trade has avoided for two years, which is what a task actually costs to complete.
Sources
6- 01Anthropic's IPO prospectus shows AI vision, surging costs | ReutersEN
- 02Anthropic IPO prospectus reveals surging costs, $42B 2025 net lossEN
- 03Sonnet 5.5 has the heaviest token use we've measured; pricing matches GPT-6 SolEN
- 04Routing LLM traffic across inference providers with TCP-style congestion controlEN
- 05Apple CEO John Ternus is reportedly planning layoffs as memory chip costs riseEN
- 06What $10k a month in sales costs you on each payment providerEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.