Inference Academy Benchmark: DeepSeek V4 Flash Costs 14.9x More on Cold Requests
Send the same prompt twice to DeepSeek V4 Flash and the second call can cost 14.9 times less. That is the widest gap in a serving benchmark inference.academy published on 7 September, covering 14 providers and three input sizes.

The measurement is not about answer quality. Inference.academy ran one controlled setup across 1k, 10k and 100k input-token targets, each paired with a 100 or 1k output-token budget. Temperature was set to zero and reasoning was switched off. Cold requests used a unique prefix. Warm requests repeated the same prompt. The site calls the results serving measurements, not task-quality scores.
Cost here means total request charges divided by input tokens, multiplied by 1M, and it includes output charges. That matters, because the headline input price on a provider page is not what the benchmark compares. The study reports cost points across 13 routes for the 10k input, 100 output-token, cold-cache condition.
The 14.9x figure is the widest gap the benchmark calls out. On DeepSeek's 100k-input, 100-output-budget requests, warm calls cost 14.9 times less than cold ones in this sample. Warm means a repeated prompt. The measured cache share decides how much input work was billed at the uncached rate.
Speed is a separate axis
Telnyx led median generation speed in all 12 conditions tested. First-token latency did not follow the same pattern. The benchmark says the fastest first token depended on the request shape. A provider that wins on time-to-first-token for a 1k prompt is not necessarily the one that wins at 100k.
"Where you send a prompt changes what you pay, how long you wait, and whether the call succeeds."
That quote is how inference.academy frames the exercise. The timing panels use cold requests with a 100-token output budget and a shared axis scale. TTFT runs from request start to the first content or reasoning delta the client receives, and it includes network and queueing time. P90 uses linear interpolation between ordered observations. Decode speed is measured after streaming starts, and it is reported as completion tokens divided by elapsed time from the first to the last output delta. Non-streaming responses get no decode-speed measurement at all.
Routing is another variable. The benchmark logged the upstream names OpenRouter returned for each request shape and cache condition. It notes that a repeated prompt may reach a different upstream, which can change its cache behavior.
For the 10k to 100 warm condition, Together accounted for 29 calls, Novita 20 and CoreWeave 17. In the 10k to 100 cold condition, Baidu appeared 15 times and Relace 13. The site is explicit that segment width is the share of calls and that the names are reported upstreams, not verified machines. It also says the distribution alone does not establish why routing changed.
Failure rates are broken out by input target and output budget, with cold and warm compared separately. The benchmark warns that a route can behave differently when the same prompt asks for a longer response. HTTP, transport and response errors all count as failures, and exact counts sit in tables below the charts. Zero means no observed failure in that sample, not a guarantee.
The practical takeaway for anyone paying inference bills is narrower than the headline. Caching changes the bill by more than an order of magnitude on long inputs, and the provider you get through a router is not fixed. The benchmark covers 14 providers and 13 routes in its cost scatter. Its own framing is the safest summary: these are serving measurements, and they describe what happened in this sample.
Sources
1All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.