DeepSeek has aimed at the strongest part of Nvidia's lead in AI: software rather than chips. On September 30 the Chinese lab open-sourced a full toolkit for programming Huawei's Ascend AI accelerators, centred on TileLang, a programming language Bloomberg described as China's answer to Nvidia's CUDA.

The release, announced on DeepSeek's WeChat account and described as "fully supported" by Huawei, ports the same low-level software DeepSeek uses to train its frontier models from Nvidia GPUs to Huawei's Ascend 950 chips. It is the most concrete step so far toward a Chinese AI stack that does not depend on American hardware or software.

What DeepSeek Released

The package contains Ascend versions of the internal libraries DeepSeek had already built for Nvidia hardware:

  • TileLang – a tile-based language for writing high-performance AI kernels, now with native code generation, automatic scheduling and synchronization for Ascend 950
  • DeepGEMM – optimised matrix multiplication, the core operation in neural networks
  • DeepEP – communication between chips at cluster scale
  • TileKernels – standard vector operations, with the same Python APIs on either an Nvidia or Ascend backend
  • FlashMLA – attention kernels for long-context workloads
  • DeepSelect – data-filtering tooling

DeepSeek says every TileLang operator used to train its V4-series models now has a high-performance Ascend version. The two companies also tuned a "supernode" cluster of 128 Ascend 950 chips, which points to large-scale training as well as inference.

Where TileLang Comes From

TileLang was not invented by DeepSeek. It was first developed by researchers at Peking University, and an Ascend version was open-sourced in September 2025. The Decoder reports that DeepSeek has used TileLang as its main tool for AGI work for about a year, first testing it on older Nvidia hardware. DeepSeek pitches the language as offering a simpler programming model than CUDA, which matters for a smaller developer community trying to write efficient code without years of GPU-specific expertise.

The new release is best seen as a large expansion of that earlier work: production-grade kernels that a frontier lab actually trains with, now published for anyone to use.

The Limits

The release does not remove every dependency. DeepGEMM-Ascend still requires Huawei hardware and CANN, Huawei's own compute software layer, and TileKernels needs either CUDA or CANN underneath. Developers moving to these tools are swapping one vendor's stack for another's, not leaving vendor stacks behind.

Hardware still matters too. Earlier reports put Huawei's Ascend 910C at roughly 60% of Nvidia H100 inference performance, and the Ascend 950 has yet to be independently benchmarked at scale. The key open question, which independent developers have not yet answered, is how much engineering effort DeepSeek's software actually saves for a team choosing Ascend over Nvidia.

Why It Matters

Nvidia's position rests on more than chip performance. An estimated four million developers write CUDA code, and years of libraries, tutorials and tooling make switching expensive. That software moat has been one of the strongest effects of US export controls: even when Chinese chips are good enough, the software around them has not been.

By publishing the kernels behind a frontier model, DeepSeek lowers that barrier for every Chinese lab and cloud provider. Research firm SemiAnalysis noted that Huawei's CANN stack was the only one besides CUDA to support DeepSeek V4 on day one, and DeepSeek had already given Chinese chipmakers, rather than Nvidia, early access to V4. The Decoder frames the release as China's AI industry closing ranks, with model makers such as Z.ai and Moonshot AI moving faster than domestic chipmakers and now pulling the hardware ecosystem along.

There are three implications for the wider industry:

  • Faster domestic substitution: Chinese labs can move training workloads to Ascend with less custom engineering.
  • A test of open-source strategy: DeepSeek is applying its open-weights approach to infrastructure, betting that adoption beats secrecy.
  • Pressure on export-control logic: if software gaps close, chip restrictions alone will slow China's AI progress less than intended.

What Comes Next

DeepSeek is reportedly planning a data centre in Inner Mongolia that could house at least 160,000 Ascend accelerators. That would be a large-scale test of whether this toolkit can support frontier training on Chinese silicon. If DeepSeek's next model is trained mostly on Ascend, September 30 may be remembered as the day the CUDA moat began to be seriously challenged.

Sources