Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

GPU Prices Climb While Cloud List Rates Fall: What Inference Actually Costs

Anthropic's IPO prospectus, reported by Reuters on 28 September, sets $518 billion of future cloud and compute obligations against a $42 billion net loss for 2025. The pricing data landing this week suggests buyers should read the second number carefully.

AI & modelsExplainerRachel NwosuPublished: 29 September 20263 min readSources 7
GPU Prices Climb While Cloud List Rates Fall: What Inference Actually Costs

The headline rates are not where the money goes. Meetrix compared AWS and GCP, checking prices on 28 September, and put an on-demand H100 at $6.88 an hour on AWS against $10.98 on GCP. That 37% gap is real. It is also smaller than the swing from running the same card at a quarter load. Meetrix's placeholder throughput of 2,500 tokens per second gives $0.76 per million tokens when the GPU is fully busy and $3.06 at 25% utilisation.

That is the whole argument in one line. Idle silicon costs more than any provider choice.

GPUAdvisor tracks 14 providers and says it last ran its data on 1 October. Its figures show the specialist clouds undercutting the hyperscalers by wide margins on older cards: an A4000 at $0.15 per GPU-hour on Hyperstack, an L40S at $0.38 on Lium, an RTX 4090 at $0.40. The site labels these as tracked approximations and tells buyers to confirm rates on the provider page. That is fair, because the same page lists an RTX 4090 at $0.74 on RunPod. Two prices, same model, nearly double.

The token bill is a different number again

Artificial Analysis published its evaluation of Anthropic's Claude Sonnet 5.5 on 29 September. The lab found the model reaches 56 on its Intelligence Index, two points behind Opus 5.5 at max effort, while using roughly 193,000 output tokens per task. That is about 60% more than Opus 5.5 and around seven times GPT-6 Astra at max effort. Anthropic kept pricing unchanged at $2 per million input tokens and $10 per million output, but the cost per task lands at $7.60, roughly 50% above Sonnet 5.

Same rate card. Different bill. The tokens do the damage.

AI Pricing Guru refreshes rates daily and last updated on 1 October. It puts the usable range for general APIs at $0.50 to $15 per million output tokens, with input typically two to eight times cheaper and cached input cutting another 75% to 90% where the provider supports it. Its calculator tracks 154 active models across more than 10 providers.

Retrieval changes the arithmetic before generation even starts. Spheron's cost breakdown, published 29 September, notes that a standard RAG setup pulling five 500-token chunks adds 2,500 tokens of context per query. A question that would cost 100 to 200 input tokens as plain chat can run 15 to 20 times higher on the input side.

Unblocked described its own routing work on 29 September, moving GLM 5.2 traffic between Baseten, Fireworks and CoreWeave on a score weighted 70% cost and 30% latency. On a fixed order favouring Baseten, that provider served 98.5% of tasks in the first full week. The company says CoreWeave's list prices ran about 45% below Baseten's, and it had no performance data to place it. That is the state of procurement: the cheapest option is often the one nobody has measured.

Comments 0

Sources

7
  1. 01Anthropic's IPO prospectus shows sweeping AI vision, surging costs: ReutersEN
  2. 02GPU Costs for LLM Inference: AWS vs GCP (2026)EN
  3. 03Cloud GPU Pricing - Cost Intelligence for H100, A100 & B200EN
  4. 04Sonnet 5.5 has the heaviest token use we've measured; pricing matches GPT-6 SolEN
  5. 05AI Token Cost Calculator 2026: 154 Models by TierEN
  6. 06RAG Inference Cost Calculator: Cost Per Token, Broken Down (2026)EN
  7. 07Routing LLM traffic across inference providers with TCP-style congestion controlEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.