Inference Cost Math: What GPU Hourly Prices Hide
GPUAdvisor's tracker, updated on 1 October, now spans 14 cloud providers selling H100, H200, A100 and B200 capacity. The spread between the cheapest and most expensive listings on the same silicon is wide enough that the headline hourly rate tells you very little.

Comparison sites have multiplied faster than the discounts. On 1 October, GPUAdvisor said its cloud GPU pricing tracker covers 14 providers and every major accelerator. Above the table sits a warning: prices are approximate, vary by region and availability, and reflect September 2026 estimates. That is the state of the market. The list price is a starting point, not a bill.
Meetrix published its own comparison on 29 September, with prices checked on 28 September for on-demand instances in US East and us-central1. Its finding is blunt. AWS is cheaper on list for an H100 or an A100; GCP is a little cheaper if a single L4 is enough. An H100 costs $6.88 an hour on AWS against $10.98 on GCP.
Utilization beats cloud choice
The same piece then undercuts its own headline. Meetrix ran the arithmetic on a placeholder throughput of 2,500 output tokens per second. Cost per million tokens moves from $0.76 to $3.06 on the AWS H100 as utilization falls from 100% to 25%. On the GCP H100 the same swing runs from $1.22 to $4.88. The article is explicit that the throughput figure is an assumption, not a measurement, and that the ratios between the rows do not depend on it.
So a 37% saving from switching clouds is worth having. It is smaller than the penalty for an idle afternoon. Meetrix also flags a data problem. Google's own GPU pricing page renders its table client-side, so the author could not read it programmatically. Three trackers agreed on $87.83 for the a3-highgpu-8g, while a fourth listed $88.49. That is a disagreement worth naming rather than averaging.
There is a structural catch on the cheap side. AWS sells the A100 only in 8-GPU shapes: p4d.24xlarge with 40 GB cards and p4de.24xlarge with 80 GB. A single A100 therefore means GCP's a2-highgpu-1g, even though each GPU costs more there. Meetrix puts the AWS 8-GPU A100 at $21.96 an hour, $2.75 per GPU, against $3.67 on the GCP single-GPU shape.
Reserved and spot pricing complicate the picture further. Meetrix reports that a 3-year reservation saves 54% to 57% on AWS, that spot discounts on H100 or A10G run 53% to 62%, and that spot on L40S is 1%. It also notes AWS cut on-demand prices by 33% for P4d and P4de and 44% for P5, measured from its 31 May 2025 prices.
One thing that did go up: EC2 Capacity Blocks for ML, where you reserve GPUs for a fixed window. The Register reported a rise of about 15% in January 2026, with p5e going from $34.61 to $39.80 an hour.
Buying the same model from three sellers
The other half of inference cost is not the GPU at all. Unblocked published an account on 29 September of routing open-weight traffic across Baseten, Fireworks and CoreWeave, all serving GLM 5.2 at different prices and speeds. Its first setup, round-robin, sent 51% of tasks to Fireworks and 49% to Baseten in its last full week, even though Fireworks charged 25% more and ran slower.
A fixed order fixed that, sending 98.5% of tasks to Baseten and 1.5% to Fireworks. Then the operational problems arrived. Changing the order needed a code change and a deploy. The circuit breaker treated five 429s in a minute as a full outage. CoreWeave's list prices, about 45% below Baseten's, sat unused because nobody had performance data. The replacement scores each provider on cost and speed, weighted 0.7 and 0.3, and moves traffic without an engineer. The post says the numbers come from its production token ledger and logs.
Model-level token behaviour feeds straight into that bill. Artificial Analysis reported on 29 September that Claude Sonnet 5.5 scores 56 on its Intelligence Index, two points behind Opus 5.5 at max effort, but uses roughly 193k output tokens per index task. It calls that the heaviest token use it has measured, about 60% above Opus 5.5 and around 7x GPT-6 Astra at max effort. Sonnet 5.5 is priced at $2 per million input and $10 per million output tokens, unchanged from Sonnet 5, yet Artificial Analysis puts it at $7.60 per task, roughly 50% higher than Sonnet 5. The company notes the evaluations ran on a pre-release deployment with a bug affecting structured outputs, fixed for the public release.
Where the money is going instead
Anthropic's IPO prospectus, seen by Reuters and reported by CNBC on 29 September, gives the demand side of this market in one line: $7.33 billion spent on compute and infrastructure in 2025, a threefold increase from 2024, more than half of $12.65 billion in total operating expenses. The same filing describes $518 billion in cloud, computing and infrastructure obligations for coming years and a net loss of $42 billion for 2025.
Memory is the constraint showing up elsewhere. Yahoo Finance reported on 30 September that Apple's new chief executive John Ternus is planning small-scale layoffs and project cancellations, according to a Bloomberg report, as memory costs rise. The article quotes former CEO Tim Cook on his final earnings call describing a 100-year flood on memory pricing. It also names SK Hynix, Samsung and Micron as having sold out much of their premium AI memory capacity through 2026.
None of this makes the per-hour GPU table wrong. It makes it incomplete. Meetrix's own framing is the honest one: fix utilization before you commit, and read down a column rather than across, because no cloud switch delivers the saving an idle afternoon costs you.
Sources
6- 01Cloud GPU Pricing — Cost Intelligence for H100, A100 & B200EN
- 02GPU Costs for LLM Inference: AWS vs GCP (2026)EN
- 03Routing LLM traffic across inference providers with TCP-style congestion controlEN
- 04Sonnet 5.5 has the heaviest token use we've measured; pricing matches GPT-6 SolEN
- 05Anthropic's IPO prospectus shows AI vision, surging costsEN
- 06New Apple CEO John Ternus is reportedly planning layoffs as memory chip costs riseEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.