Inference: a Chinese engine lifts DeepSeek V4.1 Flash throughput 6.87 times on PCIe GPUs
On eight PCIe cards with no high-speed interconnect, METASTONE's Meta-Infer engine went from 1,932 to 13,274 tokens per second by optimising kernels and communication in software, without touching the model.

The debate about AI chips is shifting. It is no longer only about who can buy the fastest hardware, but who can get the most out of the hardware already on hand. That is the argument the Chinese company METASTONE makes through its Meta-Infer inference engine, as reported by 量子位 (QbitAI).
The gap between framework and hardware
Many graphics cards can load a model, download the weights and start a service. They "work". Under load testing, though, their real performance often stays far below the advertised levels. The model is not faulty and the hardware is not broken. The loss comes from the gap between the framework and the hardware. These cards have no high-speed interconnect between them, and they do not appear in the official validation matrices of open frameworks. So operators silently fall back on generic, inefficient implementations. Communication parameters stay calibrated for other architectures, and the memory or parallelism configuration keeps default values meant for different hardware.
Six point eight seven times the throughput
Meta-Infer attacks this problem in software alone. It does not change the model structure or the semantics of the tasks. The approach has four parts: kernel completion, operator optimisation and a rewrite of communication, tuning of parallelism and memory capacity, and finally cache reuse.
The first part fixes metadata and page sizes that block the execution of native optimised operators. It replaces outdated kernels that no longer fit and widens the threshold for fast communication paths. On an environment of eight PCIe cards, with the reference community version for DeepSeek-V4.1-Flash, input throughput was only 1,932 tokens per second. After this correction phase, it rises to 5,850 tokens per second.
The second part tackles the bottlenecks. On a PCIe architecture, the bandwidth between cards is far lower than that of a dedicated interconnect, and the default strategies send time into communication waits. The engine fuses operators, overlaps computation with communication and rewrites the collective communication logic. It reaches 13,274 tokens per second, a factor of 6.87, with support for very long context.
The methodology is meant to be generic. On another model, input throughput goes from 14,546 to 22,584 tokens per second. From the co-adaptation phase alone, several models gain between 20% and 33% in throughput.
A third-party operator
METASTONE presents itself as an independent end-to-end computing operator, not tied to any model vendor. The company says it manages more than 20,000 petaflops of resources, runs about ten intelligent computing centres and two national supercomputing centres, with a commitment to availability above 99.95%. Its team claims three Gordon Bell prizes.
When hardware is constrained, software optimisation becomes the main lever for performance.
Sources
2- 01量子位 (QbitAI) : PCIe 显卡被低估了 — DeepSeek 推理吞吐翻近 7 倍ZH
- 02量子位 (QbitAI) : page d'accueil de la rédactionZH
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.