Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Anthropic's IPO maths and Apple's layoffs show what inference really costs

Anthropic's IPO prospectus, seen by Reuters and reported on 29 September, puts a $518 billion infrastructure bill next to a $42 billion net loss for 2025, while Apple's new chief executive is cutting staff as memory prices bite.

AI & modelsNewsRachel NwosuPublished: 29 September 20264 min readSources 10
Anthropic's IPO maths and Apple's layoffs show what inference really costs

Two stories landed within a day of each other. They describe the same squeeze from opposite ends. Anthropic's prospectus, seen by Reuters, commits the company to $518 billion in cloud, computing and infrastructure obligations in coming years. Apple, according to a Bloomberg report on Wednesday, is planning small-scale layoffs and project cancellations under new chief executive John Ternus.

Anthropic's numbers are in the documents Reuters reviewed: revenue grew 12-fold in 2025 to nearly $4.6 billion, the company lost more than $8 billion on an operating basis, and it spent $7.33 billion on compute and infrastructure last year, a threefold rise from 2024 and more than half of its $12.65 billion in total operating expenses. It had $20.28 billion in cash, equivalents and short-term investments as of 31 December.

The loss is not all cash.

Roughly $34 billion of the near-$42 billion net loss was an accounting charge tied to the rising estimated value of financing that could convert into Anthropic shares, not money spent running the business. Two customers produced nearly a quarter of revenue last year, and the prospectus warns that many large clients are not locked into long-term contracts.

Memory is the binding constraint

Apple's problem is more ordinary and, for buyers of AI hardware, more instructive. The company issued cautious revenue guidance in late July because it could not source enough memory chips, and margins came under pressure from higher memory prices. On his final earnings call, then-CEO Tim Cook described a "100-year flood on the memory pricing, with exponential increases in memory prices." Yahoo Finance reported that Ternus wants to avoid further price rises in 2027 if the shortage continues.

That shortage has a cause, and it is inference and training demand for high-bandwidth memory and advanced DRAM. Yahoo Finance notes that SK Hynix, Samsung Electronics and Micron have largely sold out premium AI memory capacity through much of 2026, with Nvidia, Microsoft, Amazon and Meta competing for it. Suppliers expect constrained supply into 2027.

So the cost of serving a model is not only a software question. It is a memory allocation question, and Apple is now paying for it in headcount.

What the token actually costs

Artificial Analysis measured Anthropic's Claude Sonnet 5.5 against its own index and found the model reaches 56 points, two behind Opus 5.5 at maximum effort, at identical list pricing to Sonnet 5: $2 per million input tokens and $10 per million output tokens. The catch is token use. At max effort the model consumed about 193,000 output tokens per Intelligence Index task, the heaviest the firm has measured, roughly 60% above Opus 5.5 at max and about seven times GPT-6 Astra at max. Cost per task lands near $7.60, about 50% above Sonnet 5.

Cheaper prices per token do not automatically mean cheaper tasks, which is why routing has become its own discipline. Unblocked, which runs agent workloads, described moving its GLM 5.2 traffic between Baseten, Fireworks and CoreWeave using an adaptive router that picks a provider per request on measured cost and speed. Under round-robin, Fireworks served 51% of tasks at prices 25% above Baseten's. A fixed order pushed Baseten to 98.5%. CoreWeave listed prices roughly 45% below Baseten's, and the firm had no performance data to place it.

The billing side has its own arithmetic. Flaviocopes calculated what $10,000 of monthly sales costs at 200 orders of $50: Stripe takes $350, Creem $470, Paddle $600, and Gumroad $1,450. The author discloses that Creem sponsors the site, and states every fee was checked against provider pricing pages on 29 September.

Hardware vendors are chasing the same curve

GM's battery joint venture Ultium Cells said on Tuesday that its Spring Hill, Tennessee plant will be the first in the world to mass-produce prismatic LMR cells, which it claims deliver 33% higher energy density than LFP at comparable cost. Upgrades start later this year and finish by 2028, adding 500 jobs, with the first LMR vehicle expected in 2028.

MaxLinear made a similar cost argument for networking. Its Puma 9 DOCSIS chip, announced on 29 September, supports DOCSIS 3.1, 3.1+ and 4.0 and is claimed to cut customer premises equipment costs by 30% to 50% against its predecessor, while adding Wi-Fi 8 support and DDR5 memory options as DDR4 supply tightens. Modems based on it are expected in 2027.

Two smaller engineering notes point the same way. Bitdrift published blob-stream, a Kafka alternative built to eliminate cross-availability-zone network traffic and stateless brokers, on the argument that cross-AZ costs dominate cloud streaming bills. Radim Marek benchmarked PostgreSQL 19's REPACK (CONCURRENTLY) against pg_repack and pg_squeeze on a Hetzner ccx33 and a GCE n2-standard-4, finding the online rewrite generated 18.8 GB of WAL against pg_repack's 33.5 GB on the same table.

None of this makes inference cheap. It makes the bill legible, which is a different thing.

Comments 0

Sources

10
  1. 01Anthropic's IPO prospectus shows AI vision, surging costs: ReutersEN
  2. 02New Apple CEO John Ternus is reportedly planning layoffs as memory chip costs riseEN
  3. 03Sonnet 5.5 has the heaviest token use we've measured; pricing matches GPT-6 SolEN
  4. 04Routing LLM traffic across inference providers with TCP-style congestion controlEN
  5. 05What $10k a month in sales costs you on each payment providerEN
  6. 06GM's new EV battery tech will cut costs without sacrificing performance or rangeEN
  7. 07MaxLinear claims new 'Puma 9' DOCSIS chip is a big cost-cutterEN
  8. 08Announcing blob-stream: a Kafka alternative for no fuss, low cost high volume streamingEN
  9. 09What REPACK (CONCURRENTLY) costs while it runsEN
  10. 10Anthropic IPO prospectus reveals surging costs, $42B 2025 net loss: reportEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.