Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

DeepSeek-V4.1-Flash: nearly sevenfold throughput on PCIe GPUs

A Chinese team reports a throughput jump for DeepSeek-V4.1-Flash on a PCIe system without NVLink. The case shows how much performance sits in the software.

TechnologyAnalysisGrace OkonkwoPublished: 24 September 20266 min readSources 2
DeepSeek-V4.1-Flash: nearly sevenfold throughput on PCIe GPUs

On 24 September 2026 the portal QbitAI (量子位) published a field report on METASTONE (是石科技), a Chinese company. The team ran DeepSeek-V4.1-Flash on eight GPUs wired up only over PCIe. No NVLink, so none of the fast coupling such models are usually built for.

Starting figure and final state

The authors cite a day-zero baseline of 1,932 tokens per second for input throughput. A first pass brought kernel fixes, a faster path choice for the sparse-attention-based preprocessing, swapped FP8 dense GEMM kernels and an accelerated PCIe IPC path. That alone pushed the figure to 5,850 tokens per second. More work followed: attention operator fusion, shorter drafts in speculative decoding, overlapped computation and communication, and memory parameters. The system then reached 13,274 tokens per second, roughly 6.87 times the starting value. According to the report, the context length of up to one million tokens held.

Comparative measurements show this is not a one-off. DeepSeek-V4-Flash rose from 14,546 to 22,584 tokens per second in input throughput, a factor of 1.55. The model GLM 5.3 improved from 3,236.78 to 6,222.72 tokens per second.

Why this matters for Germany

The report works on two levels. The first is technical. Models that run on hardware without a fast interconnect lose a large part of their performance advantage not to the hardware but to the memory and communication paths. Optimise those paths and you gain factors, not percentage points.

The second level concerns operations. DeepSeek-V4.1-Flash was released on 10 September 2026. It has 552 billion parameters in a mixture-of-experts layout and activates 8 billion in prefill and 16 billion in decode. The KV cache is 890 bytes per token, roughly a quarter of the value of V4-Flash. The model is under the MIT licence, processes images at up to 384 tokens per image and can be served via vLLM, SGLang or TensorRT; FP8, BF16, GPTQ, AWQ and GGUF are supported.

METASTONE states more than 20,000 petaflops of managed compute, the operation of more than ten intelligent data centres and two national supercomputing centres; three Gordon Bell Prizes are cited as a reference. For local inference, the real message of the report is the shift in value creation. Buying new hardware is not what decides. What decides is the ability to master the memory hierarchy and concurrency for a specific model.

Six times faster without new hardware: the difference between 1,932 and 13,274 tokens per second is software.

Comments 0

Sources

2
  1. 01量子位 (QbitAI): 8卡PCIe纯软优化, DeepSeek-V4.1-Flash 吞吐提升 6.87 倍ZH
  2. 02Hugging Face: DeepSeek-V4.1-Flash model card (552B MoE, MIT, 1M context)EN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.