
DeepSeek's Ascend software stack review: promising CUDA alternative, narrow audience today
Review of DeepSeek's Ascend stack: strong low-level AI tooling, but early access and validation limit production appeal.
DeepSeek's new Ascend stack is not a consumer product, and it is not a drop-in replacement for CUDA in the broad sense. It is better read as a source-available signal to AI infrastructure teams: Huawei Ascend hardware is gaining more of the low-level software that large model training and inference actually require.
The release is timely because the practical moat around Nvidia has never been only silicon. Developers need kernel authoring, matrix multiplication, expert-parallel communication, runtime compilation, tests, examples, and a path from research code to production clusters. DeepSeek and Huawei's September 30 update tackles that layer by pairing TileLang's new Ascend 950 backend with DeepSeek's Ascend ports of its compute and communication libraries.
What works on paper
The strongest part of the package is focus. DeepGEMM-Ascend targets the hot path: BF16, FP8 and FP4 matrix multiplication, MQA logits and MegaMoE operators on Ascend. DeepSeek says the package is API-compatible with DeepGEMM, which matters for teams that already know the company's Nvidia-oriented tooling. Its published README reports dense GEMM utilization up to 99.8 percent of the stated hardware limit on Ascend 950DT with CANN 9.20. That is a vendor-published result, not an independent benchmark, but it is specific enough to be useful for early technical screening.
DeepEP-Ascend is the other practical piece. It covers expert-parallel all-to-all operations for mixture-of-experts dispatch and combine, including FP8 dispatch, plus work-in-progress pipeline, bucket collective and remote-memory primitives. The published measurements show dispatch bandwidth in the 313 to 375 GB/s range across EP8 to EP128 test cases, while also noting that larger expert-parallel sizes and combine paths remain under optimization. That candor is important: this looks like a serious engineering release, but still one with sharp edges.
Where it falls short
The main limitation is access. DeepEP-Ascend's own guidance says the recommended commercial HDK and firmware baseline is planned for mid-October availability through Huawei's Atlas 850E download page, and that current performance numbers were collected on a proof-of-concept HDK with manual configuration. That makes the stack hard to recommend for production decisions today unless a team already has Huawei access and Ascend expertise.
The second limitation is ecosystem gravity. TileLang now lists official Ascend 950 support with native code generation, scheduling and synchronization, while the older tilelang-ascend adapter has examples and tutorials for A2 and A3 devices. That is helpful, but it is still a developer-facing kernel toolchain, not the mature CUDA universe of profilers, libraries, cloud availability, third-party integrations and hiring familiarity.
| Choice | Best fit | Main caveat |
|---|---|---|
| DeepSeek Ascend stack | Teams evaluating Huawei Ascend for MoE training or inference | Early public deployment path and limited independent validation |
| Nvidia CUDA ecosystem | Teams needing the broadest tooling, support and deployment options | Higher dependence on Nvidia's hardware and software platform |
Verdict
For most readers, this is a watch-list release rather than a migration trigger. It deserves attention from AI infrastructure groups under supply, sovereignty or cost pressure, especially if they already run DeepSeek-style MoE workloads or have access to Ascend 950 systems. Everyone else should wait for the public HDK, third-party benchmarks and evidence that the libraries hold up outside DeepSeek's own workloads.
Sources
Cover photo by Tima Miroshnichenko on Pexels, used under the Pexels License.
What supports the decision
Pros
- Targets real AI infrastructure bottlenecks: GEMM kernels and MoE expert communication.
- DeepGEMM-Ascend keeps API compatibility with DeepSeek's existing DeepGEMM workflow.
- Published README data gives specific Ascend 950DT performance claims to inspect.
- TileLang support makes the effort less tied to a single accelerator family.
Cons
- Recommended commercial HDK was not publicly available at the time of the release.
- Performance claims are vendor-published and need independent reproduction.
- Several DeepEP communication features are still experimental or under development.
- The stack is useful mainly to teams with Ascend hardware access and kernel expertise.
CyberOGZ Team






Comments (0)
Leave a Comment