Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Eight PCIe cards, 6.87 times the throughput: software catches up with hardware

The Chinese company METASTONE raised the input throughput of V4.1-Flash on eight PCIe cards from 1,932 to 13,274 tokens per second, without touching the model or the hardware.

TechnologyAnalysisRachel NwosuPublished: 24 September 20266 min readSources 2
Eight PCIe cards, 6.87 times the throughput: software catches up with hardware

The gap between "the model runs" and "the model delivers usable throughput" is often wider today than the gap between two generations of cards. The Chinese outlet 量子位 describes the case of 是石科技 (METASTONE), a company that does nothing but optimization, with no changes to the model structure and no change of hardware.

The starting point is the usual problem of cards with no fast interconnect between them and no place in the official support matrix of the main frameworks. The framework does not report an error. It quietly skips the accelerated paths and falls back on slow, generic implementations. On eight PCIe-only cards, in the base version from launch day, V4.1-Flash reached an input throughput of 1,932 tokens per second. The team filled in the missing kernels, fixed the fast preprocessing path for sparse attention, replaced the slow FP8 kernel for the dense GEMM operation and widened the range of fast IPC communication over PCIe. Throughput rose to 5,850 tokens per second. That is the first stage, which the authors call "coordinating model and hardware".

Second stage: communication, memory, cache

Only then does the real work begin. The engineers fused the attention and projection operators and hid computation inside the gaps of collective communication. They rewrote the communication logic for the actual PCIe bandwidth. They tuned the parallelism strategies separately for the prefill phase and the generation phase. At the end they adjusted the shares of static memory, the capacity of the KV cache and the memory reuse mechanism. Input throughput rose from 5,850 to 13,274 tokens per second, 6.87 times in total, while keeping context up to one million tokens.

The same methodology also works on other models. For DeepSeek-V4-Flash input throughput rose from 14,546 to 22,584 tokens per second (1.55 times), and for GLM5.3 from 3,237 to 6,223 (1.92 times). P95 latency of the first token fell from 141.6 seconds to 46.6 seconds, and the context limit was extended from 270,000 to 1,050,000 tokens. The first stage alone gave gains of 20 to 33 percent on several models.

What it means against top-tier hardware

The authors show the reference point honestly. At the same level of concurrency, a B300 card has about 4.9 to 5.6 times the input throughput of this eight-card machine (the comparison uses DeepSeek-V4-Flash), and for GLM-5.3 the ratio is 3.5 to 4.7 times. The distance does not disappear, but it is smaller than a comparison of specifications alone would suggest. In video generation, on the MiniMax H3 model, a 15-second video from a reference image sped up 2.48 times. The methodology was also tested on Chinese GPU cards, where sparse attention NSA had to be explicitly enabled and mechanisms written for NVIDIA and AMD platforms had to be turned off.

This is useful context for comparisons where what counts is not peak specification but the cost of producing a single token on hardware that actually sits in a server room.

Comments 0

Sources

2
  1. 01PCIe 显卡被低估了!DeepSeek 推理吞吐翻近 7 倍 (量子位)ZH
  2. 02DeepSeek V4.1 Flash — prędkość i opóźnienie (Artificial Analysis)EN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.