Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

DeepSeek V4.1 Flash: 552 billion parameters and an asymmetric encoder-decoder

Out on 10 September 2026: a multimodal Mixture-of-Experts with 552 billion parameters fires 8 billion per token on the way in and 16 billion on the way out. The architecture changed, not just the scale.

AI & modelsAnalysisRachel NwosuPublished: 26 September 20266 min readSources 3
DeepSeek V4.1 Flash: 552 billion parameters and an asymmetric encoder-decoder

On 10 September 2026 DeepSeek shipped DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model with 552 billion base parameters and a context window of up to 1 million tokens. This is no routine update. The architecture changed, and so did the way the attention cache is counted. The model card says V4.1-Flash fires just 8 billion parameters per token during prompt prefill and 16 billion during response generation, or decode.

Costs split unevenly, on purpose

The design rests on a Causal Encoder-Decoder (CED) architecture: a 40-layer Transformer split into a 20-layer causal encoder and a 20-layer decoder. In classic models every decoder layer computes its own hidden state. Here the decoder's global KV cache is projected from the encoder's final hidden states. The 8 billion to 16 billion asymmetry is deliberate. Agent workloads mostly involve reading long contexts, so the maker optimised the input side first.

The MoE layer holds 1 shared expert and 384 routed experts, 6 of which are picked per token. Then come the components the model card lists: Single-Pass mHC, a modified residual-stream mixing with the Mega-mHC kernel; Engram, a conditional memory of 196 billion parameters that token lookup queries only rarely; and DSpark, speculative decoding with semi-autoregressive draft generation and verification steered by a confidence schedule.

45 trillion tokens and vision from the first step

The model was trained from scratch on a multimodal corpus of 45 trillion tokens. Sparse attention was trained at a sequence length of 64,000 tokens. The context was extended to 1 million tokens only at the 34 trillion-token stage, after most of the training was done. After pretraining came the standard SFT recipe, then reinforcement learning, and finally on-policy distillation. The maker stresses that the changes were not in the algorithms but in the data: agentic tasks and environments were synthesised automatically at large scale.

Image input goes through the DeepSeek-ViT encoder, trained from scratch with 2D-RoPE and 3×3 pixel-unshuffle downsampling, plus a two-layer MLP projector. Images and text enter a shared space from the very start of pretraining. This is not an adapter bolted on later.

The model card on Hugging Face records 621,396 downloads in the past month and an MIT licence. Deployment variants were prepared for vLLM, SGLang, TensorRT and Docker Model Runner, along with FP8, BF16, GPTQ, AWG and GGUF quantisations. The Chinese outlet IT之家 makes the same point as the model card: in agentic scenarios, cache hits generate most of the bill, so squeezing the KV cache harder translates directly into the cost of a task. DeepSeek also announced the retirement of the older deepseek-v4-flash and deepseek-v4-flash-vision-exp, and the redirecting of requests to V4-Pro onto V4.1-Flash at its rates.

Comments 0

Sources

3
  1. 01DeepSeek-V4.1-Flash — karta modelu (Hugging Face)EN
  2. 02DeepSeek V4.1-Flash — komunikat wydaniaEN
  3. 03DeepSeek 计划 9 月 10 日前后发布 V4.1 Flash 模型 (IT之家)ZH

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.