Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

DeepSeek ships V4.1 Flash: 552 billion parameters, native multimodal, cheaper calls

DeepSeek released V4.1 Flash on September 10 Beijing time. The new MoE model carries 552 billion parameters but activates only 8 billion on input, and it cuts KV Cache demand on HBM to a quarter of what the previous generation needed.

AI & modelsNewsGrace OkonkwoPublished: 10 September 20265 min readSources 3
DeepSeek ships V4.1 Flash: 552 billion parameters, native multimodal, cheaper calls

DeepSeek released V4.1 Flash on September 10, 2026 Beijing time. The company calls it the smallest model in a new architecture series, and it handles native multimodal visual understanding. The model is already live on the DeepSeek API: point the model name at deepseek-flash and calls go through.

V4.1 Flash is a 552 billion parameter MoE model built on a new Causal-Encoder-Decoder architecture. Its input and output sides are asymmetric. The input stage activates only 8 billion parameters, the output stage 16 billion. DeepSeek says the new design was meant to raise the capability ceiling, speed up inference, increase throughput, and scale to models with larger parameter counts.

Caching is the other focus of this update. V4.1 Flash cuts KV Cache demand on HBM to a quarter of the previous generation and on SSD to an eighth. DeepSeek notes that cache hits account for a large share of the bill in Agent use cases, so compressing KV Cache can sharply lower the cost of such tasks. The company also published a long-term curve: against the first-generation model, KV Cache has shrunk 437 times.

Prices changed at the same time. During off-peak hours, input with a cache hit costs 0.02 yuan per million tokens, input with a cache miss costs 1 yuan per million tokens, and output costs 4 yuan per million tokens. Peak-hour prices are double the off-peak rates. The new pricing took effect at 12:00 on September 10, 2026.

On open source, the model weights are released under the MIT license. DeepSeek says it will fully support the open source community in building inference adaptations and will try various ways to broaden deployment. Teams with large-scale deployment needs and the resources to match can contact the team.

The fate of older versions was announced as well. V4 Flash and V4 Flash Vision Exp have been taken offline. For compatibility, the model names deepseek-v4-flash and deepseek-v4-flash-vision-exp now route to V4.1 Flash. DeepSeek also plans to retire V4 Pro in stages, after which its requests will all route to V4.1 Flash and be billed at V4.1 Flash rates. Partners WorkBuddy (including CodeBuddy) and OpenCode have fully integrated the new model.

On deployment and ecosystem, V4.1 Flash supports mainstream inference frameworks including vLLM, SGLang and TensorRT. It offers FP8, BF16, GPTQ, AWG and GGUF quantization formats and handles a context length of up to 1 million tokens. Teams in China and abroad can adopt the new model on their existing inference stacks with relatively little rework.

Comments 0

Sources

3
  1. 01DeepSeek V4.1 Flash:更强、更快、更普惠ZH
  2. 02DeepSeek-V4.1-Flash ReleaseEN
  3. 03DeepSeek V4.1 Flash 模型今日发布,V4 Pro 服务推迟到 9 月 14 日下线ZH

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.