Alibaba unveils its most powerful AI chip and opens up the software stack
At the Yunqi Conference the Chinese group announced the Zhenwu V900 chip, billed as the most powerful Chinese AI accelerator. A day later T-Head published the components of its SAIL software stack as open source.

On 22 September, at the Yunqi Conference in Hangzhou, Alibaba chief executive Wu Yongming presented the new 真武 V900 accelerator. The company calls it the most powerful Chinese AI chip on performance, with a claimed threefold edge over the previous generation M890. A day later, subsidiary T-Head published the open source work on the software stack that goes with it.
What was opened
The package, called T-Head SAIL, covers adaptation of the compute frameworks, acceleration libraries, the toolchain and communication libraries. The release lists PyTorch-for-sail, sailify, Triton-for-sail, DeepGEMM-for-sail and FlashAttention-for-sail, with PCCL and DeepEP-for-sail announced as coming. Opening the tools and not only the silicon serves one purpose: customers can adopt the chip without rewriting the models they already run.
Alibaba says more than 650 customers across more than 20 industries have taken it up, among them Ant, Xiaohongshu and the carmaker XPeng. Xiaohongshu built agents that handle workload migration and operator optimisation. The ModelScope catalogue carries 39 quantised models, including Qwen, DeepSeek and Kimi, with more than 348,000 downloads in September 2026. The competitive edge, in other words, is played as much on the developer community as on performance.
Compute stays scarce
In the background sits a supply problem. A joint report Inspur presented with IDC at AICC 2026 estimates that coverage of global demand for AI compute capacity will fall from 79% in 2024 to 71% in 2027, then rise again to 77% in 2030. The absolute gap is put at 380.9 billion dollars in 2030, almost ten times the 2024 figure. Global token consumption, meanwhile, is forecast to grow at a compound annual rate of 4,822.6% through the end of the decade, with about 4,000 trillion inference operations a year.
At the same conference Inspur presented the 元脑 SD200 Ultra system, which takes the configuration from 64 to 128 chips with expansion memory and storage of up to 8 and 64 terabytes, alongside the HC2000 platform. The competition is shifting from the single processor to the whole system: whoever integrates silicon, software and distribution scale cuts the risk of being left out of the market.
Sources
2- 01量子位 — 阿里真武 V900 发布, T-Head SAIL 开源 (Yunqi Conference, Apsara 2026)ZH
- 02量子位 — 浪潮 AICC 2026 与 IDC 《2026 中国人工智能计算力发展评估报告》ZH
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.