Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Alibaba unveils the Zhenwu V900 chip and opens up the T-Head SAIL stack

Alibaba announced the chip at the Yunqi conference on September 22 and called it the most powerful in China. A day later, T-Head added migration tools and acceleration libraries to its open-source stack.

TechnologyNewsGrace OkonkwoPublished: 25 September 20266 min readSources 2
Alibaba unveils the Zhenwu V900 chip and opens up the T-Head SAIL stack

China's cloud giants are moving from "we ship chips" to "we build the ecosystem around them together." That is how the outlet qbitai (量子位) framed it, reporting on Alibaba Group CEO Wu Yongming (吴泳铭) at the Yunqi conference on September 22. Presenting the new AI chip Zhenwu V900, he called it the most powerful computing chip in China and said it should be three times more efficient than the earlier M890.

A day later, the T-Head unit announced the next stage of opening the T-Head SAIL software stack, the counterpart of CUDA for Zhenwu chips. The open-source release includes the PyTorch-for-sail adaptation, the sailify code migration tool, the Triton-for-sail environment for writing operators, and the DeepGEMM-for-sail and FlashAttention-for-sail libraries. The first components appeared in July at the WAIC fair, when T-Head released the SDK, drivers and tools for performance analysis and debugging.

The problem it is meant to solve

The argument sounds familiar to anyone who has moved production onto a different hardware platform: migration cost matters most. Lu Shenghua, a senior director at T-Head, said plainly that the first thing customers ask is how much they will have to rewrite. A second problem shows up once a model is running. If performance drops to 30 to 40 percent of the previous level after the hardware change, nobody will accept such a migration. The third is pace. Models come out every few weeks, and customers expect a new model to run at least on its launch day.

Opening the code of the acceleration libraries shifts the balance of power for another reason too. A customer that wants to write its own operators sometimes has to reveal details of its algorithm to the chip supplier, and it does not want to do that. Access to the architecture and the code lets it do the work in-house. T-Head also describes how changes get submitted: a proposal, a code modification, automated tests and a review by maintainers, then a release. According to the company, the first contributions came both from inside Alibaba and from external customers.

Who already uses it

Figures from the presentation put Zhenwu chips at more than 650 customers across more than 20 industries. Ant Group has finished adapting its most important models for inference on Zhenwu 810E and M890 chips and is working on quantization and expert parallelism. XPeng moved training of its autonomous driving models onto a Zhenwu cluster and uses profiling tools to find bottlenecks. Xiaohongshu went further. It built its own agent for migrating and optimizing operators on top of the open SAIL code, speeding up deployment of generative recommendation models.

The ecosystem also includes a model catalog. By September 2026, 39 quantized models had been published on the ModelScope platform, covering the Qwen, DeepSeek and Kimi families, downloaded more than 348 000 times in total. The product line shows the direction: from the Hanguang 800 chip of 2019, through Yitian 710 of 2021, to today's Zhenwu series. The pressure toward openness is not altruism. Without portable code, foreign and domestic customers have no reason to switch platforms.

Comments 0

Sources

2
  1. 01量子位 (qbitai): 亮出“中国最强AI芯片”还不够,平头哥又甩出一手开源ZH
  2. 02量子位 (qbitai): 国产GPU 上的推理优化实践ZH

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.