Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

890 bytes per token: how DeepSeek shrinks the KV cache

Long context usually runs into memory, not weights. DeepSeek V4.1 Flash squeezes its KV cache to about 890 bytes per token, cutting HBM demand to a quarter of the previous generation and SSD demand to an eighth.

AI & modelsExplainerGrace OkonkwoPublished: 12 September 20265 min readSources 3
890 bytes per token: how DeepSeek shrinks the KV cache

When a model works through a long document, it keeps the attention Key and Value for every token in the context. That intermediate result is the KV cache. It grows linearly with context length, and it sets how many long sessions a single card can run at once. It is also why long context usually runs into memory rather than weights.

You can work out how much room a KV cache takes from a public model. Take Qwen3-8B and its published configuration: at BF16 or FP16, with 2 bytes per value, each token maps to 147456 bytes of KV data, about 147KB. The arithmetic is 2 for K and V, times 36 layers, times 8 KV heads, times 128 dimensions per head, times 2 bytes. That leaves out extra overhead such as cache management.

On V4.1 Flash, DeepSeek compressed it. The company's own comparison: against the previous generation, HBM demand falls to a quarter and SSD demand to an eighth. Against the first-generation model, the KV cache is already 437 times smaller. Put the official figure together with descriptions from several sources, and V4.1 Flash lands at about 890 bytes of KV cache per token.

Why does this number matter? Cache hits eat a large share of the cost in agent-style tasks. A coding agent reads code, looks things up and runs tests over and over, and the history piles up. When VRAM cannot hold it, the system clears part of the cache. The next round then has to redo Prefill for whatever it needs, and the user waits longer for the first token.

The engineering fix is tiering. Data being generated stays in HBM. What may be reused soon goes into CPU-side DDR memory. What has gone untouched longer sinks to SSD or remote storage. vLLM's KV Offloading, for instance, can offload cache blocks to CPU memory and configure a secondary storage tier, with retrieval passing through the CPU-side cache layer.

So 890 bytes is not just a number. The same VRAM can carry more concurrent long sessions, and teams running their own inference services buy fewer cards. For procurement and operations, the size of the KV cache now matters as much as the model's results, and it belongs on the selection checklist.

Comments 0

Sources

3
  1. 01DeepSeek-V4.1-Flash 模型卡EN
  2. 02DeepSeek V4.1 Flash:更少缓存,更省成本ZH
  3. 03把记忆交给 CPU,大模型会变快ZH

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.