1.5 cards for one B300: how PCIe squeezed out 6.87 times more throughput
The Chinese operator Meta-Infer pushed throughput on DeepSeek-V4.1-Flash from 1,932 to 13,274 tokens per second on eight PCIe cards. The model did not change. Only the software stack did.

The race in AI is shifting. It used to be about who buys the best hardware. Now it is about who makes better use of what they already have. The Chinese outlet qbitai (量子位) described work by 是石科技 (METASTONE). The team ran eight graphics cards connected only through a PCIe bus, with no fast interconnect between them. On DeepSeek-V4.1-Flash it got acceleration that is hard to call cosmetic.
The starting point was an ordinary community launch setup. On the same hardware, input throughput was 1,932 tokens per second. The first stage of optimization filled in missing compute kernels, swapped out a slow FP8 kernel and widened the fast communication path. Throughput rose to 5,850 tokens per second. The second stage fused attention operators, overlapped computation with communication and used separate parallelism strategies for the prefill and generation phases. That pushed the result to 13,274 tokens per second. In total, 6.87 times more, at a context of up to one million tokens.
The same method, other models
METASTONE says the same approach works more broadly. On DeepSeek-V4-Flash, input throughput rose from 14,546 to 22,584 tokens per second, or 1.55 times. With GLM5.3 the jump was from 3,237 to 6,223 tokens per second (1.92 times). P95 first-token latency fell from 141.6 to 46.6 seconds, and the supported context grew from 270,000 to 1,050,000 tokens. The company also maintains that the methodology holds up on domestic Chinese accelerators. Like the PCIe cards, those chips are not listed in the official support matrices of popular frameworks.
In the background is the operator's logistics. According to its own figures, METASTONE manages more than 20,000 P of compute power, runs more than a dozen AI data centers and two national supercomputing centers, and declares cluster availability at an SLA above 99.95 percent. The team is said to have three Gordon Bell awards to its name.
How far it still is from the flagship
An honest comparison looks less spectacular than a headline about a sevenfold gain. At the same load point, a B300 card reached about 4.9 to 5.6 times the input throughput of the eight-card PCIe machine. In image-to-video generation tasks, the advantage narrowed to 1.5 to 1.7 machines per one B300 card when LoRA was also used. Cheap cards will not catch up with top-tier hardware, but the difference stops being a gulf.
A card nobody entered into a framework's support matrix does not lose its power. It loses it to default, conservative settings.
For a reader outside China the conclusion is practical. Most self-hosted deployments are vLLM or SGLang servers on cards that are not flagships. The added value is not another card but engineering: correct kernels, communication matched to the real link throughput and different parallelism strategies for the prefill and generation phases. DeepSeek-V4.1-Flash has 552 billion parameters and a MoE architecture, and it ships under the MIT license. That makes it a good proving ground for such experiments today, especially since the model itself has a KV cache of 890 bytes per token, roughly four times smaller than V4-Flash.
Sources
3- 01量子位 (qbitai): PCIe显卡被低估了 — 内核补齐+通信重构ZH
- 02DeepSeek-V4.1-Flash – model cardEN
- 03Artificial Analysis: DeepSeek V4.1 FlashEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.