The CUDA Killer Is Free: DeepSeek Open-Sources Its Entire Huawei Ascend Toolkit
DeepSeek publicly releases TileLang, DeepGEMM, DeepEP and its full Huawei Ascend software stack — the clearest attempt yet to erase the switching cost that has locked the AI industry into Nvidia's CUDA for over a decade.
For more than a decade, the moat around Nvidia hasn’t primarily been silicon. It has been CUDA — the proprietary programming layer that every serious AI framework, kernel, and training script was built against. Switching away from Nvidia GPUs historically meant months of rewriting code, debugging, and clawing back lost performance. On September 30, 2026, DeepSeek took the most direct swing at that moat anyone has attempted: it publicly released, free to download, the entire software toolkit it built with Huawei Technologies to program Ascend AI accelerators.
What was actually released
The centerpiece is TileLang, a domain-specific programming language that DeepSeek describes as a replacement for CUDA. It is not a toy: TileLang was first tested and proven on Nvidia hardware, where DeepSeek used it to write high-performance kernels — including FlashMLA-class attention implementations — with a fraction of the code conventional CUDA requires. Roughly 80 lines of Python in TileLang can hit around 95% of hand-tuned kernel performance on some operators, and it now targets CPUs, GPUs, and NPUs alike, including Huawei’s Ascend 950.
Alongside TileLang, DeepSeek released the rest of its in-house infrastructure stack, rebuilt from scratch for Huawei silicon:
- DeepGEMM — matrix multiplication kernels for large model inference and training
- DeepEP — high-speed inter-chip communication for large-scale distributed training
- TileKernels — standard vector computation primitives
- FlashMLA — the long-context attention implementation DeepSeek famously open-sourced for its Multi-head Latent Attention architecture
- DeepSelect — data filtering for training pipelines
Together, the toolkit mirrors what DeepSeek already built for its own Nvidia GPU fleets — and mirrors are exactly the point. Any Chinese lab that wants to migrate off Nvidia now has a working, tested, free software path to Ascend hardware, maintained by a lab whose frontier models compete with the best in the world.
The technical case has been building all year
This release didn’t come out of nowhere. The groundwork has been visible for months:
- Ascend 910C delivers roughly 60% of H100 inference performance per DeepSeek’s own research reported by Tom’s Hardware, with manual optimization of the chip’s CUNN kernels pushing that figure higher.
- DeepSeek’s PyTorch library can now convert CUDA code to CUNN, Huawei’s equivalent framework — the kind of friction-removal that turns migration from theoretical to practical.
- DeepSeek and Huawei have been co-developing a supernode design linking 128 Ascend 950 chips, with as much engineering focused on how the processors talk to each other as on raw compute.
- The TileLang approach already supports most of the operators used in training DeepSeek’s V4 series — the company’s 1.6-trillion-parameter flagship released in April 2026, which matched GPT-5 and Claude Opus on major benchmarks.
That V4 detail matters more than it first appears. When DeepSeek released V4, it gave early performance-testing access to Chinese chipmakers — not to Nvidia or AMD. ByteDance, Tencent, and Alibaba subsequently rushed to lock in Ascend 950 orders once it was demonstrated that a frontier model ran well on domestic silicon. Open-sourcing the toolkit is the second half of that same play: first prove the frontier model works on Huawei chips, then hand every other Chinese lab the software to follow.
Why this is harder to counter than a chip ban
Nvidia’s China business was already thin after the H20 restrictions — the cut-down chip built specifically to comply with US export controls has faced its own licensing troubles. Bloomberg Intelligence now estimates Huawei’s Ascend line has pushed Nvidia’s China market share toward single digits, even as Chinese tech stocks broadly lag their US AI peers.
The strategic logic cuts deeper than market share. US export controls were designed under the assumption that denying chips would keep Chinese labs dependent on American suppliers. Instead, every new restriction has strengthened the incentive to build a fully sovereign stack — chips, interconnect, and now software. DeepSeek has become OpenRouter’s top model provider while running its newest infrastructure on domestic accelerators.
The uncomfortable truth for Washington is asymmetry: you can restrict a chip with a license requirement, but you cannot restrict an open-source codebase that anyone can download, fork, and improve. DeepSeek just handed away, for free, the software layer that made Nvidia genuinely hard to leave. A functioning CUDA alternative, backed by a lab with frontier-level models and a hardware partner shipping 950-series parts at scale, is a slower but far more structural threat than any single export ban.
The adoption question
None of this is automatic. Developer-ecosystem tools that mirror another company’s stack are only as valuable as the number of outside engineers who adopt them. CUDA’s advantage compounds through a decade of Stack Overflow answers, university courses, framework defaults, and pretrained kernels. TileLang and its siblings start from zero outside DeepSeek and Huawei — although the open-sourcing itself, coming from the most-watched AI lab in China, is the strongest possible seed.
What happens next depends on whether Alibaba, ByteDance, Tencent, and the next tier of Chinese AI startups standardize on this stack — or fork it into incompatible variants. But the direction of travel is now unmistakable: China’s AI infrastructure independence no longer stops at the silicon. As of today, it includes the software that programs it.
Sources
- [1] https://www.bloomberg.com/news/articles/2026-09-30/deepseek-unveils-huawei-ai-chip-tools-that-may-replace-nvidia-s
- [2] https://www.reuters.com/world/asia-pacific/deepseek-partners-with-huawei-develop-chip-programming-tools-reducing-reliance-2026-09-30/
- [3] https://startupfortune.com/deepseek-open-sources-chip-tools-that-could-let-huawei-replace-nvidia-in-china/
- [4] https://www.digitimes.com/news/a20260930VL214/deepseek-ascend-software-huawei-nvidia.html
- [5] https://github.com/tile-ai/tilelang