Six Repos Against CUDA: DeepSeek and Huawei Open-Source the Ascend Software Stack
DeepSeek has open-sourced its Ascend software stack — TileLang, DeepGEMM, DeepEP, FlashMLA, TileKernels, and DeepSelect — giving Huawei's chips a CUDA-alternative toolkit that already powers most operators in DeepSeek V4 training.
For two decades the deepest moat in accelerated computing was never the silicon — it was CUDA. Nvidia’s GPUs could be matched on paper, but the programming environment around them, the libraries, the compilers, the decade of accumulated operator code, kept developers locked in. On September 30, DeepSeek took its most direct swing yet at that moat: the lab open-sourced a six-module programming toolkit for Huawei’s Ascend AI chips, developed jointly with Huawei, that mirrors the infrastructure DeepSeek already uses in production on Nvidia hardware.
What shipped
The release, announced on DeepSeek’s official WeChat account and confirmed by Reuters, covers the entire lower layer of the AI training stack on Ascend:
- TileLang — a high-level kernel programming language and compiler toolchain, now with native Ascend 950 support including code generation, automatic scheduling, and synchronization
- DeepGEMM-Ascend — matrix multiplication library supporting BF16, FP8, and FP4, API-compatible with DeepSeek’s existing CUDA-facing DeepGEMM
- DeepEP-Ascend — the communication library that routes data to experts in mixture-of-experts models and combines their outputs during training and inference
- FlashMLA — sparse attention operators for long-context efficiency
- TileKernels — vector computation and memory-access operators for data processing
- DeepSelect — efficient data selection
Every component has a direct counterpart in the stack DeepSeek previously open-sourced for Nvidia platforms. That symmetry is the point: a developer who built against DeepSeek’s Nvidia tooling can move to Ascend while keeping familiar APIs — no rewrite required.
The TileLang bet is the story
The most strategically loaded piece is TileLang. DeepSeek describes it as offering “a simpler programming model” than CUDA, one that can “significantly improve development efficiency and simplify code logic” — while still pushing hardware to its limits, unlike other high-level alternatives that trade performance for convenience.
The proof point the lab offers is its own training runs: the majority of operators used to train the DeepSeek V4 family are now implemented in TileLang, and every TileLang operator used in DeepSeek training now has a corresponding high-performance implementation on Ascend. The Ascend version wraps Huawei’s low-level Ascend C instructions and exposes them through the same high-level interface — so the abstraction layer DeepSeek built for its models now sits, in principle, on top of either vendor’s hardware.
That is a quietly radical architecture. If your kernel layer is portable, the switching cost between Nvidia and Huawei collapses from “port your entire codebase” to “change your hardware target.” It is the software equivalent of what Huawei is doing at the supernode level.
Built on a 128-chip supernode
The engineering was not done in isolation. DeepSeek says Huawei’s team was “deeply involved” throughout development, and the two companies jointly optimized computation and inter-chip communication on a supernode built around 128 Ascend 950 chips. That matters because distributed training lives or dies on the second half of that problem: keeping thousands of accelerators fed with data so none sit idle. DeepEP-Ascend, the MoE routing layer, is precisely the component that determines whether a supernode is a real training machine or a collection of expensive, underutilized parts.
Across a number of key test cases, DeepSeek claims the compute and communication performance of these components is “already approaching the limits of the underlying hardware” — meaning the software is no longer the bottleneck on Ascend; the chips themselves are.
One caveat worth keeping in view: that claim is measured against Ascend’s ceiling, not against equivalent workloads on CUDA. A kernel at the hardware limit on one platform and a workload that costs no more than it did on the other are different claims, and no published numbers yet put a price on a full production port — engineering time, performance delta, retraining included.
Context: the ecosystem gap was the real export control
The release lands two weeks after Huawei unveiled its next-generation AI processors and supernode systems, including commercial availability of Ascend 950 clusters, and follows earlier collaboration on DeepSeek’s V4 models, which shipped with Ascend support in April. Huawei has said it expects its systems to be widely used for model training in 2027.
The unstated driver is, of course, US export controls. Chinese labs can increasingly obtain domestic accelerators with credible peak specs — but hardware parity is worthless if extracting that performance requires fighting an immature toolchain. Nvidia’s enduring advantage was never raw FLOPS; it was the twenty years of software gravity around them. By open-sourcing the exact infrastructure a frontier lab uses to train its own models, DeepSeek is attempting to bootstrap that gravity from the top down: not by building a general-purpose ecosystem and hoping frontier workloads arrive, but by shipping the frontier workload’s own foundations and inviting everyone in.
The open-source move also makes the stack auditable and forkable. Any Chinese startup — or any organization anywhere facing chip access constraints — can inspect, extend, or port these components to other accelerators. DeepSeek explicitly frames TileLang as “a useful reference for building highly usable software ecosystems around a broader range of AI chips,” which reads as an invitation well beyond Huawei.
What it means
For Nvidia, this is the erosion scenario analysts have warned about, arriving from an unexpected direction. Nobody out-engineered CUDA head-on; instead, a model lab built a portable abstraction over it, then open-sourced the port. TileLang still supports Nvidia GPUs — it was validated there first — so the language itself is not hostile to Nvidia. That is exactly what makes it dangerous: developers adopt it because it is simpler, and only later discover their code no longer cares whose chip runs it.
For Huawei, the release converts Ascend from “capable hardware with a software problem” into a platform with a frontier-lab-certified toolchain, freely available, benchmarked against hardware limits on a 128-chip system. The 2027 training promise now has visible software underneath it.
And for the broader industry, it is one more signal that the AI infrastructure stack is fragmenting along geopolitical lines — not at the model layer, where Chinese and Western systems remain interoperable, but at the substrate. The moat Nvidia built in software is now being filled in, repo by repo, in public.
All six projects are available now: TileLang and TileLang-Ascend on GitHub under tile-ai, and DeepGEMM-Ascend, DeepEP-Ascend, TileKernels, FlashMLA, and DeepSelect under deepseek-ai.
Sources
- [1] https://www.tomshardware.com/tech-industry/artificial-intelligence/deepseek-and-huawei-release-open-source-ascend-ai-programming-tools-to-reduce-reliance-on-nvidia-ecosystem-tools-include-compute-and-communication-libraries-as-well-as-ascend-support-for-tilelang
- [2] https://www.geopolitechs.org/p/deepseek-builds-for-huawei-ascend
- [3] https://github.com/tile-ai/tilelang-ascend
- [4] https://pandaily.com/deepseek-ascend-infra-oss-tilelang-deepgemm-deepep-superpod-flex
- [5] https://www.techtimes.com/articles/328396/20261002/deepseek-huawei-open-source-toolkit-lets-developers-ditch-cuda-without-rewriting-code.htm