Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

The compute race is no longer about stacking cards: two paths for domestic AI compute

At AICC 2026, an IDC report split the compute gap into two layers: the ceiling on capability and the scale of intelligence. By 2030 the shortfall could reach 380.9 billion dollars, and stacking cards alone can no longer cover the different loads of Prefill and Decode.

AI & modelsAnalysisRachel NwosuPublished: 23 September 20267 min readSources 2
The compute race is no longer about stacking cards: two paths for domestic AI compute

For years, competition in AI compute came down to one word: stack. Stack cards, stack racks, stack generators. At AICC 2026, the IDC report "2026 China Artificial Intelligence Computing Power Development Assessment Report" split the problem into two layers: the intelligence ceiling and the scale of intelligence.

Start with demand. By 2030, the report says, global AI inference tasks will grow by 4 quadrillion a year. A single multi-turn task once burned a few dozen tokens. Now it burns several hundred thousand. Multi-agent collaboration multiplies both figures. IDC data shows that from this year to 2030, global token consumption will grow at a compound rate of 4,823%.

Supply cannot keep up. The report says the global AI compute demand satisfaction rate has slid from 79% in 2024 and is expected to fall to 71% in 2027. Even if pressure on the upstream semiconductor supply chain eases, by 2030 it can only recover to around 77%. By 2030 the compute gap will reach 380.9 billion dollars, nearly ten times the 2024 figure.

Why does stacking cards no longer work? Because inference loads in the Agent era are messy. Prefill is a highly parallel, compute-dense task. Decode has to process tokens serially and leans harder on memory bandwidth. Attention frequently accesses a constantly growing KV Cache, and MoE adds sparse computation and large-scale routing and scheduling. Stacking one kind of compute tends to throw the resource mix out of balance.

Hardware vendors answer with two paths. One is capability-oriented, such as the supernode AI server Yuanbrain SD200 Ultra. The number of chips per machine rises from 64 to 128, HBM and expansion storage go up to 8TB and 64TB, and the company says a single machine can hold and run the 2.8 trillion parameter Kimi K3. It can even support deploying a 10 trillion parameter model on one machine. Unified addressing and symmetric direct connections cut All to All communication latency to 0.69 microseconds, and super operators reduce the number of operators to one tenth. The company claims more than 3x inference performance running Kimi K3.

The other is capacity-oriented, such as the multi-compute unit Yuanbrain HC2000. It uses full liquid cooling and high-flow power supply. A single cabinet delivers more than 300 kilowatts and carries up to 256 AI accelerator cards, reaching a compute density 4 times that of a traditional air-cooled cabinet, with 200Tbps of aggregated bandwidth inside the cabinet. It lets different types of chips divide work by inference stage: Prefill uses high-compute chips, Decode uses high-capacity HBM chips.

Both paths share one premise: compute has to be used more fully. The DSec paper DeepSeek published points the same way. Agent training needs to create a large number of sandboxes per second, putting far more pressure on cluster scheduling than traditional pre-training. For domestic AI compute, the next competition is not only about who has more cards, but about who can schedule heterogeneous resources more finely.

Comments 0

Sources

2
  1. 01AI 算力之争不靠堆卡!浪潮信息捅破智算能力天花板ZH
  2. 02DeepSeek 新论文公开 Agent 训练,梁文锋署名ZH

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.