Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Inference cost tools multiply as GPU prices stay opaque

A browser-based calculator published on 23 September compares inference costs across 27 models, but its own author warns the output is a planning estimate, not a vendor quote. Days later, a separate benchmark put the cost of running one model for one hour under the same spotlight.

AI & modelsAnalysisRachel NwosuPublished: 28 September 20263 min readSources 5
Inference cost tools multiply as GPU prices stay opaque

The newest entry in the inference-cost argument is not a pricing announcement. It is a calculator. Flavio Copes published an Inference Cost Calculator on 23 September. It ranks 27 models from OpenAI, Anthropic, Google, xAI, Mistral, OpenRouter and Workers AI on identical assumptions.

You set daily active users, calls per user, input and output tokens per call, and an optional prompt-cache hit rate. The tool returns a monthly bill. It also warns that the figures come from published API rates. They ignore batch pricing, enterprise discounts, image or tool surcharges, and any caching layer you built yourself.

That caveat matters more than the ranking.

The benchmark that prices an hour of thinking

On 26 September, StarSkirmish published a different kind of cost comparison. The benchmark gives each LLM one hour to write a Protoss bot in C++ for StarCraft: Brood War, then plays the bots against each other and against human-written opponents. The field is 50 LLM bots from 10 models, plus 3 demo bots and 9 competitive human written bots, for 62 entrants, run on Prime Intellect sandboxes.

GPT-6 Astra and Claude Opus 5.5 are functionally tied at the top, according to the project's own write-up, with GPT-6 Sol a clear step above the rest alongside them. The benchmark also plots score against the average API cost of one one-hour run on a log scale. In that comparison, the authors say, GPT-6 Sol offers uniquely good value.

That is a cost-per-capability framing rather than a cost-per-token one. It also admits a limit: the authors say state-of-the-art reasoning models now benefit significantly from reasoning periods longer than one hour, and that longer-running versions are planned.

Why public numbers are hard to trust

Neither of those sources gives you a GPU price. For that, the dossier has a survey published on 20 September by Nexlab, which compares self-hosted orchestrators including LocalAI, exo, GPUStack, vLLM, Ollama, Xinference and CoderAI. Its star counts were pulled from the GitHub API on 20 September. The author states plainly that where a feature could not be confirmed, the table says so instead of guessing.

The survey's useful distinctions are not about price at all. It separates whether one model can span more than one machine, whether a follow-up turn is routed to where its KV or prefix cache already sits, and whether the server can rent a GPU by itself. vLLM, at 92k stars in that snapshot, supports tensor and pipeline parallelism over Ray and prefix caching per instance. Ollama, at 181k stars, handles one machine and one model at a time. It has no cluster story beyond round-robining several URLs behind Open WebUI.

Hardware, not software, is where the money is. Brazil's PV market is not an AI story, but pv magazine reported on 26 September that average PV system prices in the country rose 7% between January and June 2026 for projects up to 300 kW, per Greener's Distributed Energy Solutions study. Kit costs for 4 kW systems rose 18.3%, from BRL 1.42/W to BRL 1.68/W.

Different commodity, same problem: the headline number moves because components move.

What the research says about the output side

One more piece of the cost picture sits on the demand side. A paper submitted to arXiv on 14 September and revised on 17 September, SlopShape, describes a 214-feature instrument that detects AI-generated commercial web content from structural signatures alone. An LLM applied it and validated it against human annotation. It reached 98.0 macro-F1 on held-out companies, and 98.1 when every AI post was reworded by its own model. The author is Jochen Madler of Sitefire.

Cheaper inference produces more text. Detecting that text is now a measurable task with published accuracy, which is its own kind of infrastructure cost.

None of these sources agree on what inference should cost, because none of them measure the same thing. The calculator measures a hypothetical bill. The benchmark measures capability per hour of API spend. The orchestrator survey measures features, not invoices. Treat any single figure as a starting assumption.

Comments 0

Sources

5
  1. 01Inference Cost CalculatorEN
  2. 02StarSkirmish BenchEN
  3. 03Self-hosted inference orchestrators compared: LocalAI, exo, GPUStack, vLLMEN
  4. 04PV system costs increase by 7% in Brazil in H1EN
  5. 05SlopShape: Identifying AI-Generated Commercial Web ContentEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.