DeepSeek V4.1 Flash: 552 billion parameters and an asymmetric encoder-decoder
Out on 10 September 2026: a multimodal Mixture-of-Experts with 552 billion parameters fires 8 billion per token on the way in and 16 billion on the way out. The architecture changed, not just the scale.

On 10 September 2026 DeepSeek shipped DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model with 552 billion base parameters and a context window of up to 1 million tokens. This is no routine update. The architecture changed, and so did the way the attention cache is counted. The model card says V4.1-Flash fires just 8 billion parameters per token during prompt prefill and 16 billion during response generation, or decode.
Costs split unevenly, on purpose
The design rests on a Causal Encoder-Decoder (CED) architecture: a 40-layer Transformer split into a 20-layer causal encoder and a 20-layer decoder. In classic models every decoder layer computes its own hidden state. Here the decoder's global KV cache is projected from the encoder's final hidden states. The 8 billion to 16 billion asymmetry is deliberate. Agent workloads mostly involve reading long contexts, so the maker optimised the input side first.
The MoE layer holds 1 shared expert and 384 routed experts, 6 of which are picked per token. Then come the components the model card lists: Single-Pass mHC, a modified residual-stream mixing with the Mega-mHC kernel; Engram, a conditional memory of 196 billion parameters that token lookup queries only rarely; and DSpark, speculative decoding with semi-autoregressive draft generation and verification steered by a confidence schedule.
45 trillion tokens and vision from the first step
The model was trained from scratch on a multimodal corpus of 45 trillion tokens. Sparse attention was trained at a sequence length of 64,000 tokens. The context was extended to 1 million tokens only at the 34 trillion-token stage, after most of the training was done. After pretraining came the standard SFT recipe, then reinforcement learning, and finally on-policy distillation. The maker stresses that the changes were not in the algorithms but in the data: agentic tasks and environments were synthesised automatically at large scale.
Image input goes through the DeepSeek-ViT encoder, trained from scratch with 2D-RoPE and 3×3 pixel-unshuffle downsampling, plus a two-layer MLP projector. Images and text enter a shared space from the very start of pretraining. This is not an adapter bolted on later.
The model card on Hugging Face records 621,396 downloads in the past month and an MIT licence. Deployment variants were prepared for vLLM, SGLang, TensorRT and Docker Model Runner, along with FP8, BF16, GPTQ, AWG and GGUF quantisations. The Chinese outlet IT之家 makes the same point as the model card: in agentic scenarios, cache hits generate most of the bill, so squeezing the KV cache harder translates directly into the cost of a task. DeepSeek also announced the retirement of the older deepseek-v4-flash and deepseek-v4-flash-vision-exp, and the redirecting of requests to V4-Pro onto V4.1-Flash at its rates.
Sources
3- 01DeepSeek-V4.1-Flash — karta modelu (Hugging Face)EN
- 02DeepSeek V4.1-Flash — komunikat wydaniaEN
- 03DeepSeek 计划 9 月 10 日前后发布 V4.1 Flash 模型 (IT之家)ZH
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.