A Model Four Times Bigger and a Chip to Match: Alibaba Bets 10 Trillion Parameters on the Machine Intelligence Era
At Apsara 2026 Eddie Wu unveiled a 5–10T-parameter Qwen successor, the Zhenwu V900 'China's most powerful' AI chip, and a 20GW data-centre buildout — Alibaba's full-stack answer to the post-Nvidia era.
Alibaba picked the morning of its annual Apsara Conference in Hangzhou to lay out the most aggressive full-stack AI bet any Chinese company has made public. Group CEO Eddie Wu told the audience that the company plans to train a new large language model with between 5 trillion and 10 trillion parameters — two to four times the size of its current 2.4-trillion-parameter flagship, Qwen 3.8 Max. In the same keynote he unveiled the Zhenwu V900, a next-generation AI chip from Alibaba’s T-Head semiconductor unit that he called the most powerful AI processor in China, and set a target of more than 20 gigawatts of global data-centre capacity by 2032. Models, silicon, and power — announced in one breath.
For a company that spent 2026 fending off questions about whether Qwen could hold its lead against DeepSeek’s open weights and Huawei-backed training pipelines, it was a statement of intent aimed at three audiences at once: Washington, which has spent two years tightening export controls on advanced accelerators; Beijing, which wants domestic champions to replace Nvidia gear; and global developers, who are deciding whose foundation models to build on next.
The model: 5–10 trillion parameters, aimed at ASI
The headline number is the parameter count. Alibaba’s current flagship, Qwen 3.8 Max, sits at roughly 2.4 trillion parameters, and Wu said the Qwen team is continuing research into model architecture and data optimisation with the goal of completing “more complex, longer-horizon tasks” and advancing toward what he called artificial superintelligence, or ASI. That is a notable rhetorical escalation: most Western labs have spent the past year soft-pedalling the AGI language while safety debates rage, while Alibaba is putting ASI on the keynote slide.
Two technical claims stand out. First, Wu said Alibaba’s proprietary M890 AI supernode — the inference cluster architecture the company has been building out — already handles inference for models above 2 trillion parameters, a capability he claimed only “a handful of companies globally” possess. Running dense inference at that scale is a genuine engineering achievement: it requires networking thousands of accelerators to behave as a single device, with the memory bandwidth and interconnect to keep them fed. Second, he said the Qwen team has made “meaningful progress” on recursive self-improvement — models that identify their own limitations, design experiments, and synthesise data to drive a cycle of self-evolution. Coming from a lab that ships open weights, that claim will draw intense scrutiny, because self-improving training loops are exactly the capability safety researchers on both sides of the Pacific have been warning about this year.
The chip: Zhenwu V900 and the 500,000-card cluster
The Zhenwu V900 is the more strategically consequential announcement. Wu said the chip delivers three times the performance of its predecessor, the M890, and that a single cluster built on it can scale to up to 500,000 cards for frontier model training and inference. He added that Alibaba expects “significant growth” in annual AI chip shipments — a signal that T-Head output is moving from internal experiment to production volume.
The context is inescapable: Chinese firms are racing to build domestic alternatives to Nvidia’s processors amid U.S. export restrictions, and every keynote this year has featured a silicon reveal. What makes the V900 different from the usual “domestic chip breakthrough” announcement is the cluster math. A 500,000-card topology is not a demonstration unit; it is the scale at which hyperscalers train frontier models. If Alibaba can actually field clusters of that size on its own silicon by the mass-production date of Q1 2027, it removes the single biggest external constraint on the 10-trillion-parameter ambition announced on the same stage. Alibaba Cloud also said it will begin bringing its AI supernodes online at commercial scale this quarter.
The power: 20GW by 2032
The third leg is energy. Wu set a target for Alibaba Cloud’s global data-centre capacity to surpass 20 gigawatts by 2032, extending a previously announced $53 billion three-year AI infrastructure spend. He described customer demand for AI as “exceptionally robust” and said it is accelerating Alibaba Cloud’s revenue growth — but acknowledged that global shortages across the AI data-centre supply chain are limiting how fast the company can expand. “The industry’s mid-to-long-term demand far outpaces our supply capabilities,” he said.
That admission matters as much as the target. Transformer racks, high-bandwidth memory, liquid cooling, and grid interconnects are now the binding constraints on AI buildouts everywhere, and Alibaba is telling investors that the bottleneck is physical, not financial. The 20GW figure puts Alibaba in the same infrastructure weight class as the largest Western hyperscalers, and it arrives the same week Nvidia formalised its DSX Ready program certifying batteries and cooling gear for AI factories — evidence that the whole industry is converging on power and thermals as the next battlefield.
The philosophy: “Machine Intelligence” as Industrial Revolution
Wu framed the moment in sweeping terms, describing the dawn of an era of “Machine Intelligence” comparable to the Industrial Revolution. He predicted that machines would eventually produce more than 1,000 times the “thinking” of all humanity, up from less than 3% today, and offered a historical analogy: AI coding, he said, is like the light bulb of the electrical age — an early application, not the breakthrough product. “The truly groundbreaking products of the Machine Intelligence era have not yet arrived.”
It is a familiar rhetorical move — anyone who has watched Sam Altman or Jensen Huang hold a stage will recognise the genre — but the substance behind it differs. Alibaba is not merely promising smarter chatbots; it is vertically integrating the entire stack, from the instruction set of its own accelerators through the model weights it open-sources to the gigawatts it intends to control. That is a bet that the next decade of AI progress will be won by whoever owns the physical substrate, not just the algorithms.
Why it matters
Three takeaways for anyone tracking the frontier:
- Scale race resumes. After a year in which labs publicly debated slowing down frontier training, Alibaba just announced the largest planned model build in its history. The “pace the frontier” discourse dominating U.S. headlines has no obvious purchase in Hangzhou.
- Silicon sovereignty is arriving. A 500,000-card domestic cluster with 3× predecessor performance, mass production in Q1 2027, is the strongest public signal yet that China’s leading cloud vendor believes it can train frontier models without Nvidia.
- Open weights plus vertical integration. Qwen remains the most widely deployed open-weight family in the world, and newly appointed Qwen head Dayiheng Liu — promoted just yesterday, days before this keynote — now inherits both the 5–10T mandate and the recursive self-improvement agenda.
The uncomfortable question left hanging in the Hangzhou air: if machines are to think 1,000× humanity’s output, who is pacing that frontier? A U.S. federal court is currently entertaining exactly that question about four American labs. Alibaba’s answer, delivered in a single keynote, was speed — on every layer of the stack at once.
Sources
- [1] https://www.reuters.com/business/retail-consumer/alibaba-plans-ai-model-with-5-trillion-10-trillion-parameters-unveils-new-chip-2026-09-22/
- [2] https://economictimes.indiatimes.com/tech/artificial-intelligence/alibaba-plans-ai-model-with-5-trillion-to-10-trillion-parameters-unveils-new-chip/articleshow/134401408.cms
- [3] https://aiweekly.co/ai-news-today