890 bytes per token: how DeepSeek shrinks the KV cache
Long context usually runs into memory, not weights. DeepSeek V4.1 Flash squeezes its KV cache to about 890 bytes per token, cutting HBM demand to a quarter of the previous generation and SSD demand to an eighth.

When a model works through a long document, it keeps the attention Key and Value for every token in the context. That intermediate result is the KV cache. It grows linearly with context length, and it sets how many long sessions a single card can run at once. It is also why long context usually runs into memory rather than weights.
You can work out how much room a KV cache takes from a public model. Take Qwen3-8B and its published configuration: at BF16 or FP16, with 2 bytes per value, each token maps to 147456 bytes of KV data, about 147KB. The arithmetic is 2 for K and V, times 36 layers, times 8 KV heads, times 128 dimensions per head, times 2 bytes. That leaves out extra overhead such as cache management.
On V4.1 Flash, DeepSeek compressed it. The company's own comparison: against the previous generation, HBM demand falls to a quarter and SSD demand to an eighth. Against the first-generation model, the KV cache is already 437 times smaller. Put the official figure together with descriptions from several sources, and V4.1 Flash lands at about 890 bytes of KV cache per token.
Why does this number matter? Cache hits eat a large share of the cost in agent-style tasks. A coding agent reads code, looks things up and runs tests over and over, and the history piles up. When VRAM cannot hold it, the system clears part of the cache. The next round then has to redo Prefill for whatever it needs, and the user waits longer for the first token.
The engineering fix is tiering. Data being generated stays in HBM. What may be reused soon goes into CPU-side DDR memory. What has gone untouched longer sinks to SSD or remote storage. vLLM's KV Offloading, for instance, can offload cache blocks to CPU memory and configure a secondary storage tier, with retrieval passing through the CPU-side cache layer.
So 890 bytes is not just a number. The same VRAM can carry more concurrent long sessions, and teams running their own inference services buy fewer cards. For procurement and operations, the size of the KV cache now matters as much as the model's results, and it belongs on the selection checklist.
Sources
3All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.