← All posts / Industry

Fifteen Thousand Chips, One Machine: Huawei Unveils the Atlas 960 SuperPoD and Pulls Its Ascend Roadmap Forward

At Huawei Connect 2026 in Shanghai, Huawei shipped the Atlas 960 SuperPoD — 15,488 Ascend NPUs across 220 cabinets fused by UnifiedBus and new near-packaged optics — and announced that its Ascend 960DT training chip will arrive three quarters early, in Q1 2027.

Fifteen Thousand Chips, One Machine: Huawei Unveils the Atlas 960 SuperPoD and Pulls Its Ascend Roadmap Forward

On Thursday morning in Shanghai, Huawei rotating and acting chairman David Wang Tao walked onto the stage at Huawei Connect 2026 and made the strongest statement yet about where China’s AI compute industry is heading. The company unveiled the Atlas 960 SuperPoD — a computing cluster of up to 15,488 Ascend NPUs arranged across 220 cabinets and 2,200 square metres — alongside an upgraded version of its UnifiedBus interconnect and, for the first time, near-packaged optics (NPO) built directly into the SuperPoD architecture.

The subtext was barely subtext. “Supernodes are an inevitable choice for building ultra-large-scale data centres,” Wang said. And in case anyone missed who the audience for that message was, he added: “Huawei is set to begin mass production of the world’s first NPO product with an integrated light source.”

What Huawei actually announced

The Atlas 960 SuperPoD is the successor to the Atlas 950 SuperPoD, which Huawei rolled out commercially last year and which has pushed total deployment of the Atlas SuperPoD line past 1,000 units. Where the Atlas 950 clusters 8,192 Ascend NPUs, the Atlas 960 scales to 15,488 Ascend cards. Huawei claims a 2.3-fold improvement in training and a 2.5-fold improvement in inference for 10-trillion-parameter models, compared with its predecessor.

Three technology layers make the scale-up possible:

Near-packaged optics. NPOs are a placement structure that moves optical interconnect closer to the compute package, enabling higher speeds at lower energy use. Huawei says this is the first time NPOs have been introduced into its SuperPoD architecture, and that it will begin mass production of the world’s first NPO product with an integrated light source.

UnifiedBus plus Hi-ONE. The combination of Huawei’s interconnect protocol and its Hi-ONE optical system allows a SuperPoD to connect as many as 4,000 processors, extending high-speed optical connections deeper into the computing system. UnifiedBus — known as Lingqu in Chinese — is Huawei’s direct answer to Nvidia’s NVLink, the proprietary interconnect that binds GPUs within a cabinet into a coherent whole. Huawei also unveiled a portfolio of 11 chips powered by UnifiedBus spanning computing, storage, management and interconnect for SuperPoDs and Superclusters.

The “one machine” thesis. A Huawei paper posted to arXiv on the same day, authored by Huawei researcher Liao Heng, describes a 256K-node-class single AI computing system already being deployed — extending the same architecture far beyond an individual SuperPoD. The paper’s framing is striking: the system is described as “one machine,” not a network of more than 200,000 separate machines, with UnifiedBus and Huawei’s Nested BP computing model spanning the system end-to-end.

The roadmap just got faster

Perhaps the most consequential news for the competitive landscape is the schedule. Wang said development of the Ascend 960DT — the AI chip designed for model training — is ahead of schedule: “Our Ascend 960 chip is progressing faster than expected, with performance doubling. We have brought the target launch for Ascend 960 DT forward by three quarters, and it will be ready in the first quarter of 2027.”

The inference-oriented Ascend 960 PR chip moves up a quarter as well, landing in Q3 2027. Beyond that, Huawei committed to an annual upgrade cadence: “In 2028 and 2029, we will successively launch the Ascend 970 and 980. Guided by innovation based on the Tau Law, we aim to keep doubling computing performance, while also significantly increasing bandwidth, memory capacity and interconnect bandwidth.”

That is an aggressive public commitment for a company operating under US semiconductor sanctions, and it leans on the Tau Scaling Law and LogicFolding architecture Huawei unveiled in May — semiconductor breakthroughs the company says can deliver transistor performance equivalent to a 1.4-nanometre process node by 2031 without the advanced lithography tools that export controls keep out of China. Huawei has said it plans to deploy LogicFolding on its flagship Ascend 990 chip around 2030.

Why this matters beyond China

The announcements land in a week when the AI infrastructure conversation is dominated by two themes: the staggering power budgets of frontier training runs, and the geopolitics of who is allowed to build them. Huawei’s SuperPoD will face off directly against US rivals in the global computing arena — most notably Nvidia’s NVL rack-scale systems — while UnifiedBus challenges NVLink at the protocol layer. The customers for both are increasingly the same: governments, sovereign-cloud programmes, and hyperscalers outside the US-China duopoly who want optionality.

The timing is also pointed. The announcements came ahead of high-level trade talks between President Xi Jinping and US counterpart Donald Trump later this month, where AI and semiconductor exports are expected to top the technology agenda. Every benchmark Huawei can put on the board before those talks strengthens Beijing’s hand that export controls are accelerating, rather than containing, China’s domestic capability.

There are still real headwinds. Huawei’s Ascend stack historically trailed Nvidia on software maturity — the CUDA ecosystem remains the industry’s default — and HBM supply for domestic chips has been a persistent constraint. A 15,488-card SuperPoD is also an enormous capital commitment in a market where utilisation economics are still being worked out. But the direction of travel is unmistakable: in 2021 roughly 90 percent of China’s AI chips came from foreign suppliers, and today Huawei is shipping thousand-unit runs of supernodes, pulling chip roadmaps forward by three quarters, and publishing papers about quarter-million-node “single machines.”

For the global AI industry, the Shanghai keynote is a reminder that the compute layer of the AI race is now genuinely multipolar. Nvidia’s grip on the frontier remains strong, but the assumption that US export controls would freeze China a generation behind is looking weaker with each Huawei Connect.