Inference cost tools multiply as GPU prices stay opaque
A browser-based calculator published on 23 September compares inference costs across 27 models, but its own author warns the output is a planning estimate, not a vendor quote. Days later, a separate benchmark put the cost of running one model for one hour under the same spotlight.

The newest entry in the inference-cost argument is not a pricing announcement. It is a calculator. Flavio Copes published an Inference Cost Calculator on 23 September. It ranks 27 models from OpenAI, Anthropic, Google, xAI, Mistral, OpenRouter and Workers AI on identical assumptions.
You set daily active users, calls per user, input and output tokens per call, and an optional prompt-cache hit rate. The tool returns a monthly bill. It also warns that the figures come from published API rates. They ignore batch pricing, enterprise discounts, image or tool surcharges, and any caching layer you built yourself.
That caveat matters more than the ranking.
The benchmark that prices an hour of thinking
On 26 September, StarSkirmish published a different kind of cost comparison. The benchmark gives each LLM one hour to write a Protoss bot in C++ for StarCraft: Brood War, then plays the bots against each other and against human-written opponents. The field is 50 LLM bots from 10 models, plus 3 demo bots and 9 competitive human written bots, for 62 entrants, run on Prime Intellect sandboxes.
GPT-6 Astra and Claude Opus 5.5 are functionally tied at the top, according to the project's own write-up, with GPT-6 Sol a clear step above the rest alongside them. The benchmark also plots score against the average API cost of one one-hour run on a log scale. In that comparison, the authors say, GPT-6 Sol offers uniquely good value.
That is a cost-per-capability framing rather than a cost-per-token one. It also admits a limit: the authors say state-of-the-art reasoning models now benefit significantly from reasoning periods longer than one hour, and that longer-running versions are planned.
Why public numbers are hard to trust
Neither of those sources gives you a GPU price. For that, the dossier has a survey published on 20 September by Nexlab, which compares self-hosted orchestrators including LocalAI, exo, GPUStack, vLLM, Ollama, Xinference and CoderAI. Its star counts were pulled from the GitHub API on 20 September. The author states plainly that where a feature could not be confirmed, the table says so instead of guessing.
The survey's useful distinctions are not about price at all. It separates whether one model can span more than one machine, whether a follow-up turn is routed to where its KV or prefix cache already sits, and whether the server can rent a GPU by itself. vLLM, at 92k stars in that snapshot, supports tensor and pipeline parallelism over Ray and prefix caching per instance. Ollama, at 181k stars, handles one machine and one model at a time. It has no cluster story beyond round-robining several URLs behind Open WebUI.
Hardware, not software, is where the money is. Brazil's PV market is not an AI story, but pv magazine reported on 26 September that average PV system prices in the country rose 7% between January and June 2026 for projects up to 300 kW, per Greener's Distributed Energy Solutions study. Kit costs for 4 kW systems rose 18.3%, from BRL 1.42/W to BRL 1.68/W.
Different commodity, same problem: the headline number moves because components move.
What the research says about the output side
One more piece of the cost picture sits on the demand side. A paper submitted to arXiv on 14 September and revised on 17 September, SlopShape, describes a 214-feature instrument that detects AI-generated commercial web content from structural signatures alone. An LLM applied it and validated it against human annotation. It reached 98.0 macro-F1 on held-out companies, and 98.1 when every AI post was reworded by its own model. The author is Jochen Madler of Sitefire.
Cheaper inference produces more text. Detecting that text is now a measurable task with published accuracy, which is its own kind of infrastructure cost.
None of these sources agree on what inference should cost, because none of them measure the same thing. The calculator measures a hypothetical bill. The benchmark measures capability per hour of API spend. The orchestrator survey measures features, not invoices. Treat any single figure as a starting assumption.
Sources
5- 01Inference Cost CalculatorEN
- 02StarSkirmish BenchEN
- 03Self-hosted inference orchestrators compared: LocalAI, exo, GPUStack, vLLMEN
- 04PV system costs increase by 7% in Brazil in H1EN
- 05SlopShape: Identifying AI-Generated Commercial Web ContentEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.