Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

KV cache at 890 bytes per token: how DeepSeek squeezed attention memory

The model card lists 890 bytes of cache per token, about a quarter of what V4-Flash uses, and HBM demand at a quarter of the previous generation.

AI & modelsExplainerRachel NwosuPublished: 26 September 20265 min readSources 4
KV cache at 890 bytes per token: how DeepSeek squeezed attention memory

The most interesting number in the DeepSeek-V4.1-Flash spec is not about parameters. It is about memory. The model card says the global KV cache takes 890 bytes per token, about a quarter of what it takes in DeepSeek-V4-Flash. The vendor's note describes the same effect from the infrastructure side: HBM demand drops to a quarter and SSD demand to an eighth against the previous generation.

That matters because agentic tasks have the model read long contexts over and over. Cache hits make up a large part of the bill, and the KV cache still has to sit somewhere physical. Once a context runs to hundreds of thousands of tokens, the cache stops being an implementation detail. It becomes the main limit on throughput and on what a server costs to run.

Three mechanisms instead of one trick

The team got there with several independent methods. The first is SWA Bounded Replay. It rebuilds missing attention-window states by replaying only the last tokens of the window, so they never have to be written to SSD for good. The persistent KV footprint falls to about an eighth of what it was before.

The second is Comprossed Sparse Attention 2 (CSA2). It puts every attention layer into one of three modes: Full, Reindex or Reuse. Layers share the main KV and the indexer key, and sparse attention indices get reused. In the decoder, the hierarchical sparse indexer keeps the deeper indexing layers inside a candidate pool built by the first layer in Full mode. Indexing cost stops growing with context length.

The third is storing the main KV in FP4: values in E2M1 with one E4M3 scale per 16 channels. Only all three together produce 890 bytes per token.

What it changes in practice

A smaller KV footprint means more parallel sessions and longer contexts fit on the same server, and less data has to move between memory and disk. The vendor's note promises lower serving costs and says the savings go to users, with peak and off-peak pricing to match. DeepSeek calls it "pushing the limits of KV cache compression": a race not over who has more parameters but over who keeps long-context memory cheaper.

The Chinese portal IT之家 points out that on the deployment side V4.1-Flash has gone into, among others, a national supercomputer network, where cutting HBM and SSD demand is what makes it possible to serve many users at once.

Comments 0

Sources

4
  1. 01DeepSeek-V4.1-Flash — karta modelu (Hugging Face)EN
  2. 02DeepSeek V4.1-Flash — komunikat wydaniaEN
  3. 03DeepSeek V4.1 Flash 模型上线国家超算互联网 (IT之家)ZH
  4. 04DeepSeek V4.1 Flash — dane wydania i benchmarki (Artificial Analysis)EN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.