← All posts / Meta

160,000 Ascend Chips and No Nvidia in Sight: Inside DeepSeek's Inner Mongolia Mega-Cluster

Bloomberg reports DeepSeek will deploy at least 160,000 Huawei Ascend 950DT accelerators at a 1 GW data center in Inner Mongolia — the largest known all-domestic AI cluster and a stress test of China's push to replace Nvidia.

160,000 Ascend Chips and No Nvidia in Sight: Inside DeepSeek's Inner Mongolia Mega-Cluster

For most of the past two years, the phrase “Chinese AI cluster” carried an asterisk: the racks might be stamped Huawei, but somewhere in the supply chain lurked Nvidia — smuggled, stockpiled, or laundered through shell companies. A Bloomberg report from September 4 may mark the moment the asterisk comes off. DeepSeek, the Hangzhou lab that kicked off the open-weights revolution, plans to deploy at least 160,000 of Huawei’s next-generation Ascend 950DT accelerators at a massive data center under construction in Inner Mongolia — according to people familiar with the matter, it would be one of the largest known AI clusters ever built with virtually no American silicon inside.

The project is a statement of intent on three levels at once: technical, industrial, and geopolitical. And its fate will tell us a great deal about whether China’s two-year sprint toward compute self-sufficiency is actually working.

What we know

The cluster is designed around the Ascend 950DT, Huawei’s flagship rack-scale accelerator announced at HC 2025 in September 2025. Each chip pairs Huawei’s HiZQ 2.0 high-bandwidth memory — 144 GB per accelerator with 4 TB/s of bandwidth — with the UnifiedBus 2.0 fabric that binds thousands of chips into a single logical machine. In Huawei’s own supernode topology, 8,192 Ascend 950DTs occupy roughly 160 cabinets and deliver on the order of 8 EFLOPS of AI compute; an Atlas 950 SuperCluster chains 64 supernodes into a system of more than 520,000 chips.

At 160,000 chips, DeepSeek’s buildout is roughly 20 Atlas 950 supernodes’ worth of silicon — enough capacity, on paper, to train frontier-scale models without a single H-class GPU. Reports indicate the site is scoped for around 1 GW of power, which would place it among the largest single-tenant AI campuses anywhere in the world, and comfortably in the tier of xAI’s Colossus or Meta’s Ohio complexes.

The choice of Inner Mongolia is not incidental. The region has spent a decade marketing itself as China’s data center frontier: cool, dry air for free cooling, coal-adjacent grids supplemented by some of the country’s best wind and solar resources, and land that provincial officials are eager to allocate. DeepSeek already runs training and inference workloads in the corridor around Ulanqab, alongside ByteDance and Alibaba. The new mega-cluster extends that blueprint from “cheap place to rent racks” to “sovereign AI infrastructure at gigawatt scale.”

Why it matters: the software is the hard part

The instinct is to treat this as a hardware story — a chip-count headline, a Nvidia-lost-sales narrative. It isn’t. The binding constraint on Chinese domestic AI compute has never really been FLOPS; it has been the software stack. Huawei’s Cann toolkit, MindSpore framework, and Ascend-compatible PyTorch forks have historically lagged CUDA by a wide margin in maturity, developer mindshare, and raw distributed-training throughput.

That is precisely what makes DeepSeek the right tenant for this experiment. The lab’s entire identity is software extremism — squeezing frontier capability out of constrained hardware. DeepSeek’s V-series and R-series releases demonstrated that careful architecture choices (multi-head latent attention, sparse MoE routing, distillation pipelines) could deliver frontier-adjacent performance on a fraction of the training budget its American rivals burned. Its engineers then open-sourced every detail, which is why “DeepSeek tricks” now appear in the training recipes of labs on three continents.

Running that playbook on Ascend silicon is a different order of challenge — it requires re-instrumenting the whole stack, from kernel-level operator libraries to fault tolerance across a fabric that cannot lean on InfiniBand or Nvidia’s NCCL. But if any lab treats CUDA-moat pessimism as a solvable engineering problem rather than a law of nature, it is this one. Chinese media have reported for months that DeepSeek was hiring aggressively for chip-software roles; the Inner Mongolia cluster is where those hires either pay off or don’t.

The supply-side question

Even granted the software, 160,000 accelerators is a staggering ask of Huawei’s foundry situation. Ascend 950-series chips are produced by SMIC on its N+2 process, which lacks access to EUV lithography. Output yields are widely understood to be far below TSMC’s, and Huawei has been forced to shard production across multiple fabs and packaging lines. The “at least” in Bloomberg’s phrasing does a lot of work: whether DeepSeek actually receives all 160,000 units on schedule — and how many are fully binned 950DTs versus cut-down variants — depends on SMIC’s ability to keep Climbing the yield curve through 2026 and 2027.

This is the quiet race underneath the loud one. American export controls were designed on the theory that constraining lithography would constrain model capability over time. The Inner Mongolia cluster is the most direct test yet of that theory: if Huawei can ship the volume, and DeepSeek can tame the software, the compute gap between Chinese and American frontier labs narrows from “existential” to merely “expensive.”

Context: the money is already moving

The cluster does not exist in a financial vacuum. DeepSeek closed its first external round in June 2026 — more than 50 billion yuan (~$7.4 billion) at a valuation north of $50 billion, with state-backed investors and an unusual structure that kept founder Liang Wenfeng firmly in control. The lab has since been reported to be preparing a STAR Market IPO filing as early as this year, at valuations discussed in the $71–74 billion range. A gigawatt-scale, all-domestic compute base is exactly the kind of asset that story needs: it converts DeepSeek from “brilliant but infrastructure-dependent” to “brilliant and sovereign,” which matters both to Beijing’s priorities and to public-market investors who remember what chip sanctions did to previous Chinese AI champions.

It also compounds. Every DeepSeek improvement to the Ascend software stack — every operator kernel, every communication optimization, every fault-recovery path — accrues to Huawei’s ecosystem as a whole, lowering the barrier for every other Chinese lab that follows. The cluster is a single tenant’s buildout, but its externality is national.

What to watch

  • Deployment cadence. Watch for shell-company procurement records and Huawei supply-chain disclosures that reveal whether chip deliveries are tracking the 160,000 figure or slipping against it.
  • Training signals. If DeepSeek’s next flagship model ships with a technical report noting Ascend-native training runs — even partial ones — that is the ballgame for the “CUDA moat is permanent” thesis.
  • SMIC yields. Any public datapoint on N+2 yield improvements translates almost linearly into how fast this cluster and its successors fill up.
  • The Nvidia counterfactual. Jensen Huang declared just this week that “AGI has arrived” on the back of Blackwell demand. The Inner Mongolia buildout is the other half of that sentence — the arrival of a compute bloc that does not intend to buy a single one of his chips to get there.

The bottom line

The headline number is 160,000 chips. The real number to watch is smaller and harder to see: the fraction of DeepSeek’s frontier training compute that runs on domestic silicon a year from now. If that fraction is high, September 2026 will be remembered not as the month of one big cluster, but as the month the bifurcation of global AI infrastructure stopped being a policy forecast and became a physical fact on the Mongolian steppe.