Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Inference Pricing Splinters: Cache Hits, Voice Latency and Self-Hosting

DeepSeek V4 Flash requests with 100k input tokens cost 14.9 times less warm than cold in one controlled sample, according to a benchmark published by inference.academy on 7 September 2026.

AI & modelsAnalysisRachel NwosuPublished: 27 September 20265 min readSources 3
Inference Pricing Splinters: Cache Hits, Voice Latency and Self-Hosting

That single ratio, measured across 14 routes, captures what has happened to inference pricing this year. The headline rate per million tokens no longer decides a bill. Cache state, request shape and which upstream answers the call now move cost by an order of magnitude.

The inference.academy setup was deliberately narrow. It used 1k, 10k and 100k input-token targets, each with a 100 or 1k output-token budget, temperature 0 and reasoning off. Cold requests carried a unique prefix; warm requests repeated the same prompt. The author called these serving measurements, not task-quality scores. Cost was total request charges divided by input tokens, multiplied by 1M, including output charges, so it is not the listed input-token price.

Cache share is the sum of cached input tokens divided by the sum of input tokens for successes reporting both. It is a token share, not the percentage of requests that hit cache. The warm sequence includes its initial request.

Routing added a second variable the study could not fully explain. On 1k input and 100 output tokens with a cold cache, OpenRouter returned 20 upstream names, with Baidu at 18 calls, Relace at 14 and AkashML at 13. The same request shape warm returned 14 names, led by Relace at 24, CoreWeave at 16 and DeepSeek at 11. The report notes that a repeated prompt may reach a different upstream, which can change its cache behaviour. It also states plainly that the distribution alone does not establish why routing changed.

Telnyx led median generation speed in 12 of 12 conditions, according to the same benchmark. The fastest first token depended on the request shape. A model that looks quick on a short prompt may not hold that position on a long one.

Voice models compete on milliseconds, then on price

Latency benchmarks in speech have followed a similar pattern. Nari Labs wrote on 14 September 2026 that, as of mid September, it tops both the Coval speech-to-text and text-to-speech benchmarks. It ranks first on latency and second on word error rate for STT, and second on latency and first on WER for TTS. The company adds a caveat worth keeping: Coval's figures can fluctuate every 30 minutes.

Its Qwen3-ASR Fast endpoint is ranked first in time-to-final-segment at a p50 of 44 ms, with a WER of 3.6%. That places it second behind AssemblyAI's Universal 3.5 Pro at 3.5%. Nari says the Fast endpoint at $0.12 per hour ties for the second-lowest price among models with known public rates in Coval's pricing directory. Universal 3.5 Pro costs 3.75 times more, it says, while Deepgram Nova 3 costs 2.4 times more. The Standard endpoint would be cheapest at $0.06 per hour.

On the synthesis side, Nari's Qwen3-TTS Fast model is ranked second in time-to-first-audio at a p50 of 63 ms with a WER of 3.8%, which the company says is first on quality. Only vui from Fluxions posts a lower median TTFA, at 49 ms, from a 300M parameter model compared with the 1.7B Qwen3-TTS Nari serves.

At $10 per 1M characters, Nari says its Fast endpoint is tied for cheapest on Coval's directory. ElevenLabs Eleven v3 Conversational costs 5x more and Cartesia Sonic 3.6 costs 6.5x more.

The sharpest comparison in the post is between endpoints serving the same weights. Nari reports that the official Qwen3 TTS Flash Realtime endpoint sits at 8.8% WER and a median TTFA in the hundreds of milliseconds. Baseten's dedicated Qwen3-TTS endpoint records 6.0% WER and a median TTFA in the low hundreds of milliseconds. Same model, different serving stack, different numbers.

Self-hosting has a price too, and it is in operations

Nexlab's 20 September 2026 survey of self-hosted orchestrators covers LocalAI, exo, GPUStack, Xinference, Ollama, vLLM and CoderAI, with GitHub star counts pulled from the API that day. The useful column is not stars but what each tool does across machines.

  • Ollama, at 181k stars, is one machine and one model at a time, with no cluster story beyond round-robining several Ollama URLs.
  • vLLM, at 92k stars, handles tensor and pipeline parallelism over Ray but does not manage models, users or placement.
  • LocalAI, at 49k stars, added a distributed mode in June 2026 with libp2p discovery and a router aware of VRAM and prefix caches; the survey says its raw LLM throughput trails a dedicated engine by some tens of percent.
  • exo, at 47k stars, scales on Apple Silicon with MLX and RDMA over Thunderbolt 5, reporting 3.2x on four devices, but is CPU-only on Linux in September 2026.
  • GPUStack, at 5.7k stars, offers users, keys, metering and Prometheus, and supports nine accelerator vendors.

LiteLLM, at 59k stars, routes between endpoints and clouds but runs no model. The survey says that distinction is often missed.

Across all three sources, the pattern is the same: cost has stopped being a single published rate. A gateway that ignores cache state, a voice endpoint that ignores tail latency or an orchestrator that ignores failure rates will each produce a bill or a user experience that the sticker price did not predict.

What none of these sources settle is how durable the gaps are. Cache discounts depend on routing, routing changes between calls, and benchmark positions shift within the hour. Buyers comparing two providers on a single number are comparing the wrong thing.

Comments 0

Sources

3
  1. 01DeepSeek V4 Flash across 14 providers: cost, speed and cachingEN
  2. 02Show HN: Nari Qwen3-TTS and Qwen3-ASREN
  3. 03Self-hosted inference orchestrators compared: LocalAI, exo, GPUStack, vLLMEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.