Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Inference: a Chinese engine lifts DeepSeek V4.1 Flash throughput 6.87 times on PCIe GPUs

On eight PCIe cards with no high-speed interconnect, METASTONE's Meta-Infer engine went from 1,932 to 13,274 tokens per second by optimising kernels and communication in software, without touching the model.

TechnologyAnalysisRachel NwosuPublished: 24 September 20266 min readSources 2
Inference: a Chinese engine lifts DeepSeek V4.1 Flash throughput 6.87 times on PCIe GPUs

The debate about AI chips is shifting. It is no longer only about who can buy the fastest hardware, but who can get the most out of the hardware already on hand. That is the argument the Chinese company METASTONE makes through its Meta-Infer inference engine, as reported by 量子位 (QbitAI).

The gap between framework and hardware

Many graphics cards can load a model, download the weights and start a service. They "work". Under load testing, though, their real performance often stays far below the advertised levels. The model is not faulty and the hardware is not broken. The loss comes from the gap between the framework and the hardware. These cards have no high-speed interconnect between them, and they do not appear in the official validation matrices of open frameworks. So operators silently fall back on generic, inefficient implementations. Communication parameters stay calibrated for other architectures, and the memory or parallelism configuration keeps default values meant for different hardware.

Six point eight seven times the throughput

Meta-Infer attacks this problem in software alone. It does not change the model structure or the semantics of the tasks. The approach has four parts: kernel completion, operator optimisation and a rewrite of communication, tuning of parallelism and memory capacity, and finally cache reuse.

The first part fixes metadata and page sizes that block the execution of native optimised operators. It replaces outdated kernels that no longer fit and widens the threshold for fast communication paths. On an environment of eight PCIe cards, with the reference community version for DeepSeek-V4.1-Flash, input throughput was only 1,932 tokens per second. After this correction phase, it rises to 5,850 tokens per second.

The second part tackles the bottlenecks. On a PCIe architecture, the bandwidth between cards is far lower than that of a dedicated interconnect, and the default strategies send time into communication waits. The engine fuses operators, overlaps computation with communication and rewrites the collective communication logic. It reaches 13,274 tokens per second, a factor of 6.87, with support for very long context.

The methodology is meant to be generic. On another model, input throughput goes from 14,546 to 22,584 tokens per second. From the co-adaptation phase alone, several models gain between 20% and 33% in throughput.

A third-party operator

METASTONE presents itself as an independent end-to-end computing operator, not tied to any model vendor. The company says it manages more than 20,000 petaflops of resources, runs about ten intelligent computing centres and two national supercomputing centres, with a commitment to availability above 99.95%. Its team claims three Gordon Bell prizes.

When hardware is constrained, software optimisation becomes the main lever for performance.

Comments 0

Sources

2
  1. 01量子位 (QbitAI) : PCIe 显卡被低估了 — DeepSeek 推理吞吐翻近 7 倍ZH
  2. 02量子位 (QbitAI) : page d'accueil de la rédactionZH

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.