Chinese compute suppliers say software is what counts
"One and a half 6000D cards beat one B300," a Chinese cloud compute supplier sums up its inference optimisation. Inspur points to a compute gap worth 381 billion dollars by 2030.

Talk about AI costs usually starts with the price of graphics cards. Chinese infrastructure supplier 是石科技 (METASTONE) argues the bigger reserves sit in software. The company tuned how it runs the DeepSeek-V4.1-Flash model so that it works on eight cards wired only through PCIe, with no fast interprocessor bus.
量子位 described the result: input throughput rose from 1 932 to 13 274 tokens per second, a factor of 6.87. The team filled out the compute kernel, rebuilt how the cards talk to each other and reused the context cache. The company sums it up by saying that "one and a half 6000D cards beat one B300". Cheaper, slower hardware wins against more expensive kit if the processing is arranged well. The firm runs more than 20 000 P of compute across several data centres and declares service availability above 99.95 percent.
The second voice comes from Inspur (浪潮信息), which based its AICC 2026 conference talk on IDC data. According to those figures, global demand for AI compute will be met at 79 percent in 2024, 71 percent in 2027 and around 77 percent in 2030. The gap between needs and supply will reach 380.9 billion dollars, ten times more than in 2024. Token consumption is expected to grow more than 48 times year on year on a cumulative basis, and inference tasks are to reach 4 000 trillion per year in 2030.
Inspur answers with hardware. The SD200 Ultra system holds 128 chips, 8 TB of memory and 64 TB of storage in a single machine, which is meant to let models with a trillion parameters be served without splitting them across many racks. All-to-all communication latency has been cut to 0.69 microseconds.
For buyers of compute, the price per token is no longer a function of hardware prices alone. More and more often it comes down to how efficiently a supplier uses what it already has, and that changes the rules of tenders for cloud services.
The conclusion for buyers is practical. A several-fold rise in throughput came from rewriting cache handling and spreading work across the available chips, so a decision to buy a new generation of hardware should be preceded by a software audit. A supplier that can show measured values on the same infrastructure is now selling savings, not compute. It is also the simplest way Chinese data centre operators answer export restrictions: where the most powerful cards cannot be bought, what counts above all is the efficiency of what already stands in the server room.
Sources
2All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.