Next upSF Pitch Night by the AI Collective - #SFTechWeek
News

DeepSeek ports DeepGEMM kernels to Huawei Ascend 950

DeepSeek’s MIT-licensed DeepGEMM-Ascend gives developers an open-source kernel path for Huawei Ascend 950 NPUs, while its compatibility and performance claims remain independently untested.

D
Sep 30, 2026 · 2 min read

DeepSeek released DeepGEMM-Ascend on September 30. The MIT-licensed port brings its matrix-multiplication kernel library to Huawei Ascend 950 neural processing units. Developers get code and a build path for BF16, FP8 and FP4 general matrix multiplication, or GEMM, along with MQA logits and MegaMoE operations.

The release expands DeepSeek’s developer software lineup after its September V4.1 Flash model release.

DeepSeek says the Ascend package is fully API-compatible with its upstream DeepGEMM project, with the same package name, interfaces and development workflow available on other supported platforms. That compatibility statement has not been independently tested. The initial release is documented for Ascend 950 devices, and support for other Ascend generations is not established.

The project requires a Huawei Ascend NPU, CANN 9.20, torch_npu, Python 3.10 or later, and compiler support for the C++20 format library. To set it up, the repository tells developers to clone the project with its submodules, run develop.sh to link the required headers and build the C++ extension, then install the package with pip without build isolation.

At the kernel level, DeepSeek says the library wraps Ascend matrix multiply-add primitives to abstract fractal layouts, alignment constraints, address calculations and other low-level details. The implementation uses Ascend-specific sparse data loading and coroutine-based pipelining. The repository also describes a difference from Nvidia hardware: pairs of UE8M0 scaling factors along the K dimension are packed into int16 values and stored in MN-major order.

In DeepSeek’s own tests, dense GEMM reached as much as 99.8% of the stated hardware limit, including 431 TFLOPS against a listed 432-TFLOPS limit for one BF16 case. The company says it measured the kernels on an Ascend 950DT with CANN 9.20, using bench_msprof with a cold L2 cache and cases drawn from the DeepGEMM test suite. Those results have not been independently reproduced.

DeepSeek credits Huawei with technical support and engineering expertise during development.

More news