Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

DeepSeek V4 Flash costs 14.9x less on a warm cache, benchmark of 14 providers finds

Sending the same prompt twice to DeepSeek V4 Flash can cut request cost by 14.9 times, according to a serving benchmark published by inference.academy on 7 September that routed the model across 14 providers.

AI & modelsNewsGrace OkonkwoPublished: 28 September 20263 min readSources 1
DeepSeek V4 Flash costs 14.9x less on a warm cache, benchmark of 14 providers finds

The benchmark measured one controlled setup: 1k, 10k and 100k input-token targets, each paired with a 100 or 1k output-token budget, temperature 0 and reasoning off. Cold requests carried a unique prefix. Warm requests repeated the same prompt. The site is explicit that these are serving measurements, not task-quality scores.

Cost in the study is not the listed input-token price. It is total request charges divided by input tokens and multiplied by one million, so output charges are folded in. Across the sample, DeepSeek V4 Flash's 100k-input, 100-output-budget requests cost 14.9 times less warm than cold.

Routing is not stable between calls

The benchmark also logged which upstream names OpenRouter returned for each request shape. Those names shift with the cache condition. For 1k-input, 100-output cold calls the list was led by Baidu with 18 calls, Relace with 14 and AkashML with 13. The same shape warm was led by Relace with 24, CoreWeave with 16 and DeepSeek with 11.

At 100k input and a 100-token output budget, cold requests went mostly to Baidu (7), DigitalOcean (3), Relace (3) and Wafer (3). Warm requests of the same shape went to Relace (14), Together (6) and CoreWeave (5). The authors caution that segment width is only the share of calls, that names are reported upstreams rather than verified machines, and that this distribution alone does not establish why routing changed.

On speed, Telnyx led median generation speed in all 12 conditions. First-token latency depended on the request shape rather than on a single winner. The site defines time to first token as running from request start to the first content or reasoning delta received by the client, including network and queueing time. Decode speed is reported completion tokens divided by elapsed time from the first to the last output delta. Non-streaming responses produce no decode-speed measurement.

Timing includes successful calls only.

Cache share and failure rates

The study reports cache share as the sum of cached input tokens divided by the sum of input tokens for successes that reported both, measured on repeated 10k prompts with a 100-token output budget. That is a token share, not the percentage of requests that hit a cache, and the warm sequence includes its initial request.

Failure rates are broken out by input target, output budget and cache condition. The authors note a route can behave differently when the same prompt asks for a longer response. Failure rate is failed calls divided by measured calls in that condition, counting HTTP, transport and response errors. Exact counts and descriptive intervals sit in tables below the charts, and the raw records are downloadable. A stated zero means no observed failure in the sample. The benchmark marks unmeasured fields as not reported rather than zero.

For buyers, the practical reading is narrow. The gap between cold and warm is large enough that prompt reuse, not provider choice alone, may dominate the bill for repeated workloads. But the same numbers show that a repeated prompt can land on a different upstream, which changes whether a cache discount applies at all.

Comments 0

Sources

1
  1. 01DeepSeek V4 Flash across 14 providers: cost, speed and cachingEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.