ResearchHot

DeepSeek Open-Sources Its AI Stack for Huawei Ascend Chips

DeepSeek has open-sourced a full set of software tools that let its AI models run on Huawei's Ascend chips. The components mirror the ones it already released for Nvidia hardware, and for developers working on domestic AI stacks, they remove a long-standing roadblock.

Evan BrooksEvan Brooks
Heat: 1,450
An editorial illustration of server chips in a data center rack with abstract software overlays

DeepSeek has open-sourced a full set of software tools that let its AI models run on Huawei's Ascend chips. The components mirror the ones it already released for Nvidia hardware, and for developers working on domestic AI stacks, they remove a long-standing roadblock.

DeepSeek released the tools on September 30, 2026, publishing a package of infrastructure components built for Huawei's Ascend computing platform. The release covers a programming language and compiler, computing libraries, and distributed communication libraries, and each piece lines up with an equivalent that DeepSeek had already open-sourced for Nvidia GPUs.

The move matters for anyone who writes software that runs on AI chips. Nvidia's CUDA software layer has been the default for years, and building an alternative takes more than hardware. It takes tools that developers can actually use to get performance out of that hardware.

What DeepSeek open-sourced for Huawei Ascend

The headline component is TileLang, a high-level programming language and compiler. DeepSeek positions it as a simpler route to hardware performance than writing raw low-level code, comparing it to CUDA in purpose but aiming for less complexity. The Ascend version wraps the chip's low-level instructions, so developers get a higher-level way to write code without giving up performance.

Alongside TileLang, the release bundles five more components, each mapped to a role in the training and serving pipeline.

Here's what each one does.

Component

What it does

TileLang

High-level language and compiler for writing chip operators

DeepGEMM

Speeds up general matrix operations, the core math in neural networks

DeepEP

Handles communication across many devices running together

TileKernels

Provides standard vector computations and memory access routines

FlashMLA

Runs sparse attention for long-context models

DeepSelect

Filters data efficiently during processing

Every component has a counterpart on the Nvidia side, which makes the intent clear. DeepSeek built these tools for its own training runs and is now offering the Ascend versions to everyone.

That last point matters.

Why TileLang and the Ascend stack matter

AI models run on thousands of chips at once, and the software that coordinates them decides how much of the hardware's theoretical speed you actually get. That layer is where Nvidia's lead is hardest to close, because it's built on years of developer habit as much as on engineering.

TileLang was first proven on Nvidia hardware. DeepSeek says most of the operators used to train its V4 series models were built with it, and the Ascend version now carries every TileLang operator that those training runs needed, each with a working high-performance implementation.

That "every operator has a match" claim is the important part. A tool that covers 80% of the operations a model needs is a curiosity. One that covers all of them is something a team can build on without falling back to hand-written code.

Coverage is the whole game.

DeepSeek's Ascend performance numbers and what they mean

DeepSeek reports that its computing and communication components have reached close to the hardware's limits in several key tests. The company cites figures like 99.8% of the hardware limit for DeepGEMM's matrix operations and 98% for a mixed-expert setup called MegaMoE.

Those are developer claims, not results verified by an outside party, so treat them as targets the team is reporting rather than settled benchmarks. The practical reading is that the tools are tuned to squeeze most of the available performance out of the chips, which is the bar any production AI team cares about.

DeepSeek also says the work was done with Huawei's support, and that the two teams optimized a 128-chip supernode design based on the Ascend 950. A supernode is a cluster of chips wired to behave like one larger machine, which is how modern training runs scale up.

What the DeepSeek Ascend release means for developers

For teams building on Huawei hardware, the release cuts down on the custom engineering that used to sit between a model and the chip. Instead of writing device-specific code for every operator, developers can use a higher-level tool that already has working implementations for the operations models rely on.

Why does that matter so much?

The bigger question is whether this helps Ascend chips get used more broadly. Software tooling is often the reason teams pick one chip family over another, and a fuller stack lowers the cost of switching. This release doesn't settle that contest, but it removes one of the practical excuses for staying put.

There's no pricing here, because the components are open source. The repositories are hosted on GitHub under DeepSeek's organization, and the TileKernels library ships under the MIT license, one of the most permissive open-source terms available. That makes the code easy to adopt, modify, and ship in commercial products.

For an outside observer, the clearest signal is where the effort went. DeepSeek didn't just port a model. It rebuilt the plumbing underneath the model, the kind of work that pays off slowly and quietly rather than in a launch demo.

What to watch after DeepSeek's Ascend release

The first thing to track is independent testing. If outside teams reproduce the near-hardware-limit figures, the toolkit earns credibility beyond DeepSeek's own reports. If results vary widely across workloads, the picture is more limited.

The second is adoption. Open-sourcing a stack is the start of the work, not the end. Watch whether other model makers and cloud providers build on these components, because that's what turns a one-company project into an actual alternative.

The tools are available now in DeepSeek's GitHub repositories. Developers who want to test them can clone the components and run the benchmarks against their own workloads, which is the fairest way to judge what the numbers really mean.

Share This Story

Sources

Related AI News

Your AI Looks Great in Demos. LLM Evaluation Tools Prove Whether It Works
Research

Your AI Looks Great in Demos. LLM Evaluation Tools Prove Whether It Works

LLM evaluation tools catch what demos hide. This guide covers platforms, open-source frameworks and benchmarks, explains how LLM-as-judge scoring works, and shows how to start testing cheaply.

Heat: 1,210
What Does LLM Stand For? Large Language Models, Explained Plainly
Research

What Does LLM Stand For? Large Language Models, Explained Plainly

LLM stands for large language model. This plain-English guide breaks down what the letters mean, how these models turn your prompt into text one token at a time, and where their limits show up in daily use.

Heat: 1,280
Why Good Models Break in Production: AI Deployment Challenges, Solved
Research

Why Good Models Break in Production: AI Deployment Challenges, Solved

A model that works in a notebook can still fail in production. This explainer covers the AI deployment challenges that matter, from latency and cost to quality drift, plus rollouts and monitoring that keep a model working.

Heat: 1,050
RAG in 2026: Why Your AI Assistant Keeps Citing Sources
Research

RAG in 2026: Why Your AI Assistant Keeps Citing Sources

RAG, short for retrieval augmentation, lets a language model answer questions about data it never trained on by looking it up first. This explainer covers how the pipeline works and why citations aren't a guarantee.

Heat: 1,080