AWS holds the cheaper H100, but utilisation decides your inference bill
AWS lists an H100 at $6.88 an hour against $10.98 on Google Cloud, a 37% gap that matters far less than the fourfold swing between a busy GPU and an idle one, according to pricing checked on 28 September.

AWS lists an H100 at $6.88 an hour against $10.98 on Google Cloud, a 37% gap that matters far less than the fourfold swing between a busy GPU and an idle one, according to pricing checked on 28 September.
That comparison comes from meetrix.io, which says it checked on-demand list prices on 28 September for US East on AWS and us-central1 on Google Cloud. The same page states plainly that an H100 costs four times more per token at 25% utilisation than at full load, on either cloud. The author also warns that the tokens-per-second figure in the worked example is a placeholder, not a benchmark.
The table, and who is cheapest in it
Meetrix puts AWS about 37% cheaper per H100 and 25% cheaper per A100 than Google Cloud on demand. The one place Google wins is the small L4 card: g2-standard-4 comes in at about $0.70 an hour against $0.805 for AWS's g6.xlarge, a difference the article itself calls roughly $75 a month.
The A100 row hides a catch. AWS only sells the A100 in eight-GPU shapes, p4d.24xlarge with 40GB cards and p4de.24xlarge with 80GB, at $2.75 per GPU-hour and $21.96 for the node. If you need a single A100, Google's a2-highgpu-1g at $3.67 an hour is the practical option even though each GPU costs more.
On sourcing, meetrix is unusually candid. It says AWS figures come from Vantage's EC2 tracker, updated 28 September 2026, and that Google's own GPU pricing page renders its table client-side so it could not be read programmatically. The GCP numbers are therefore ones price trackers publish, and three of them agreed on $87.83 for the a3-highgpu-8g while a fourth listed $88.49. That is a small disagreement, but it is a disagreement, and it is worth knowing which side of it your budget sits on.
Utilisation beats shopping around
The arithmetic in the meetrix post uses 2,500 output tokens per second on one H100, which the author labels an assumption rather than a measurement. At 100% busy, AWS works out at $0.76 per million tokens and GCP at $1.22. At 50% busy those become $1.53 and $2.44. At 25% busy, $3.06 and $4.88.
Read down a column, not across, the post advises.
It is the right instruction. Switching clouds saves 37% on the same H100. Letting the card sit idle three quarters of the time costs you four times as much. One is a procurement decision, the other is a scheduling one, and only the second compounds daily.
GPUAdvisor, which says it tracks 14 providers and last ran its daily check on 1 October 2026, shows where that shopping-around logic leads. Its listings put Lium's L40S at $0.38 per GPU-hour, flagged as 89% below AWS, Thunder Compute's RTX A6000 at $0.35, and Hyperstack's A4000 at $0.15. The same page notes that GPU-specialist clouds offer two to four times lower per-GPU pricing than hyperscalers, and that prices are approximate and vary by region. Specialists are not the same product as a hyperscaler, but the spread is not noise either.
Commitment discounts move the number too, and they cut both ways. Meetrix reports a three-year AWS reservation saving 54% to 57%, with the caveat that GPUs age quickly. GPUAdvisor puts one-year commitments at 35% to 40% for always-on inference endpoints and GCP Spot at up to 70% off A100 and H100 instances. Meetrix adds that Spot on AWS offers 53% to 62% off for H100 and A10G today, but only 1% for L40S.
The capacity product that went the other way
Not every price fell. Meetrix cites The Register's report of a rise of about 15% in January 2026 on EC2 Capacity Blocks for ML, with p5e going from $34.61 to $39.80 an hour. It notes that this is a separate product from the on-demand rates in the comparison table, which is exactly the kind of distinction that gets lost when teams quote a single GPU number in a planning meeting.
The context for all of this is a supply squeeze that reaches well beyond GPUs. Yahoo Finance reported on 30 September that new Apple CEO John Ternus is planning small-scale layoffs and project cancellations, citing a Bloomberg report, as memory costs rise. The piece quotes former CEO Tim Cook on his final earnings call describing a 100-year flood on memory pricing, and notes that SK Hynix, Samsung and Micron have largely sold out premium AI memory capacity through much of 2026. The same buyers are bidding for compute and memory.
Anthropic's IPO prospectus, seen by Reuters and reported on 29 September, puts a number on that appetite: $7.33 billion spent on compute and infrastructure in 2025, a threefold surge from 2024, accounting for more than half of $12.65 billion in total operating expenses, against $518 billion in cloud, computing and infrastructure obligations planned for coming years. That is one buyer's forward commitment, and it is larger than most clouds' annual GPU revenue.
Meanwhile the software side keeps inflating the bill. Artificial Analysis reported on 29 September that Claude Sonnet 5.5 uses roughly 193,000 output tokens per Intelligence Index task at max effort, the heaviest token use it says it has measured, around 7x GPT-6 Astra at max. Anthropic priced it identically to Sonnet 5 at $2/$10 per million input/output tokens, so the cost lands in the token count, not the rate card, and Artificial Analysis puts it at $7.60 per task, about 50% higher than Sonnet 5.
Routing is the other lever, and it is getting automated. Unblocked described on 29 September an adaptive router that scores providers on cost and speed, weighted 0.7 and 0.3, and moves traffic when one degrades. Before it, round-robin sent 51% of GLM 5.2 tasks to Fireworks at prices 25% above Baseten. After a fixed order, Baseten served 98.5%. The company says a third provider, CoreWeave, listed prices about 45% below Baseten's, and it had no performance data to place it.
None of this makes the AWS versus GCP question wrong. It makes it the second question. Measure the utilisation first, then argue about the hourly rate.
Sources
6- 01GPU Costs for LLM Inference: AWS vs GCP (2026)EN
- 02Cloud GPU Pricing — Cost Intelligence for H100, A100 & B200EN
- 03New Apple CEO John Ternus is reportedly planning layoffs as memory chip costs riseEN
- 04Anthropic's IPO prospectus shows sweeping AI vision, surging costs: ReutersEN
- 05Sonnet 5.5 has the heaviest token use we've measured; pricing matches GPT-6 SolEN
- 06Routing LLM traffic across inference providers with TCP-style congestion controlEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.