Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Chinese compute suppliers say software is what counts

"One and a half 6000D cards beat one B300," a Chinese cloud compute supplier sums up its inference optimisation. Inspur points to a compute gap worth 381 billion dollars by 2030.

BusinessAnalysisDr. Amara PatelPublished: 26 September 20265 min readSources 2
Chinese compute suppliers say software is what counts

Talk about AI costs usually starts with the price of graphics cards. Chinese infrastructure supplier 是石科技 (METASTONE) argues the bigger reserves sit in software. The company tuned how it runs the DeepSeek-V4.1-Flash model so that it works on eight cards wired only through PCIe, with no fast interprocessor bus.

量子位 described the result: input throughput rose from 1 932 to 13 274 tokens per second, a factor of 6.87. The team filled out the compute kernel, rebuilt how the cards talk to each other and reused the context cache. The company sums it up by saying that "one and a half 6000D cards beat one B300". Cheaper, slower hardware wins against more expensive kit if the processing is arranged well. The firm runs more than 20 000 P of compute across several data centres and declares service availability above 99.95 percent.

The second voice comes from Inspur (浪潮信息), which based its AICC 2026 conference talk on IDC data. According to those figures, global demand for AI compute will be met at 79 percent in 2024, 71 percent in 2027 and around 77 percent in 2030. The gap between needs and supply will reach 380.9 billion dollars, ten times more than in 2024. Token consumption is expected to grow more than 48 times year on year on a cumulative basis, and inference tasks are to reach 4 000 trillion per year in 2030.

Inspur answers with hardware. The SD200 Ultra system holds 128 chips, 8 TB of memory and 64 TB of storage in a single machine, which is meant to let models with a trillion parameters be served without splitting them across many racks. All-to-all communication latency has been cut to 0.69 microseconds.

For buyers of compute, the price per token is no longer a function of hardware prices alone. More and more often it comes down to how efficiently a supplier uses what it already has, and that changes the rules of tenders for cloud services.

The conclusion for buyers is practical. A several-fold rise in throughput came from rewriting cache handling and spreading work across the available chips, so a decision to buy a new generation of hardware should be preceded by a software audit. A supplier that can show measured values on the same infrastructure is now selling savings, not compute. It is also the simplest way Chinese data centre operators answer export restrictions: where the most powerful cards cannot be bought, what counts above all is the efficiency of what already stands in the server room.

Comments 0

Sources

2
  1. 01PCIe显卡被低估了!内核补齐+通信重构,DeepSeek推理吞吐翻近7倍ZH
  2. 02AI算力之争不靠堆卡!浪潮信息捅破智算能力天花板ZH

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Dr. Amara Patel

Dr. Amara Patel

Economy, business and world

Dr. Amara Patel covers business, world affairs and the economy for FLASH24, working from filings, central bank statements and trade data rather than press releases, and she does not let company spin stand in for numbers. She checks revenue recognition, debt covenants and currency effects line by line against audited reports and regulatory disclosures. Her week includes calls with analysts, logistics operators and trade lawyers, and she watches the calendar for rate decisions, earnings dates and port and freight updates, comparing each against prior quarters. Outside the desk she tracks tech-company accounts and rides cargo bikes, which keeps her close to both the balance sheets she reads and the supply chains she covers. She does not publish a figure she cannot trace to a primary document.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.