DeepSeek open-source tools are giving Huawei’s Ascend AI chips a stronger software foundation as Chinese technology companies intensify efforts to reduce their dependence on Nvidia’s CUDA ecosystem.
On September 30, 2026, DeepSeek released six software modules adapted for Huawei’s Ascend platform. The collection includes tools for AI computation, communication between accelerators and performance optimization, with an Ascend-compatible version of DeepSeek’s TileLang programming language among the most significant additions.
The announcement matters because building a serious alternative to Nvidia requires more than producing a fast AI processor.
Developers also need programming languages, optimized libraries, communication frameworks and tools that make it relatively straightforward to turn theoretical chip performance into useful computing power.
That software layer is where Nvidia has enjoyed one of its greatest advantages for years.
DeepSeek and Huawei are now trying to narrow that gap.
DeepSeek Open-Source Tools Bring Six Components to Ascend
DeepSeek has released six core infrastructure components for Huawei’s Ascend computing platform.
They include TileLang, DeepGEMM, DeepEP, TileKernels, FlashMLA and DeepSelect.
These components mirror software that DeepSeek had previously developed around Nvidia-based computing infrastructure, but they have now been adapted to work with Huawei hardware.
The tools address different parts of running demanding AI workloads.
Some are designed to improve calculations carried out on accelerators.
Others handle communication between chips or optimize commonly used operations in large AI models.
Taken together, they help create more of the infrastructure developers need to train and run advanced artificial intelligence models using Huawei processors.
The fact that DeepSeek is making the tools open source is also important.
Other developers can inspect the code, experiment with it and potentially contribute improvements instead of depending entirely on proprietary software controlled by one company.
TileLang Is the Most Direct Challenge to Nvidia CUDA
Among the DeepSeek open-source tools, TileLang stands out because it addresses one of the most difficult parts of competing with Nvidia: programming AI accelerators efficiently.
TileLang is a high-level programming language designed for writing high-performance computing kernels.
A kernel is a specialized piece of software responsible for performing a particular calculation efficiently on processors such as GPUs or AI accelerators.
DeepSeek says the new version of TileLang supports Huawei’s Ascend 950 accelerators and includes native code generation, automatic scheduling and synchronization.
The idea is to allow developers to express complex computing operations without manually managing every low-level hardware detail.
That puts TileLang in the same broad problem space that has helped make Nvidia CUDA so influential.
It does not mean TileLang already matches CUDA in ecosystem size, maturity or developer adoption.
But it gives Huawei an additional programming layer through which developers can potentially extract more performance from Ascend hardware.
Why Nvidia CUDA Is So Difficult to Challenge
Nvidia’s dominance in artificial intelligence is often described as a hardware story.
Its GPUs are used extensively for training and running large AI models.
But the company’s competitive position also comes from software.
CUDA was introduced in 2006 and has since grown into a broad computing ecosystem containing programming tools, optimized mathematical libraries, debugging utilities and support across major AI frameworks.
That ecosystem has accumulated years of documentation, community knowledge and developer expertise.
As a result, replacing Nvidia hardware can create another problem.
Software designed around CUDA may need to be changed, optimized or rewritten to run efficiently elsewhere.
That switching cost has helped reinforce Nvidia’s position.
DeepSeek and Huawei are therefore addressing a problem that cannot be solved by semiconductor development alone.
If Huawei wants Ascend chips to become more widely used for advanced AI, developers need an increasingly capable software environment around them.
DeepSeek Open-Source Tools Focus on the Software Gap
This is why the DeepSeek announcement is strategically important.
Instead of focusing exclusively on benchmark numbers, the companies are building infrastructure developers can actually use.
DeepSeek described its broader goal as creating a programming environment that is high-level, easy to use and capable of approaching the full performance of the underlying hardware.
That is effectively a software ecosystem problem.
A chip can deliver impressive theoretical performance, but if developers find it difficult to program or optimize, much of that capacity may remain underused.
Mature tools reduce that friction.
They allow engineers to spend more time building AI systems and less time manually adapting software to hardware.
That has been one of CUDA’s strongest advantages.
DeepGEMM Targets a Core AI Calculation
One of the other tools is DeepGEMM.
GEMM stands for general matrix multiplication.
Although the terminology sounds technical, matrix multiplication is one of the fundamental calculations underlying modern artificial intelligence.
Large language models repeatedly perform enormous numbers of matrix operations during both training and inference.
Making those calculations faster can therefore improve overall AI performance.
DeepGEMM is designed to provide highly optimized matrix multiplication capabilities.
Bringing that software to Huawei’s Ascend platform means developers can use a DeepSeek-developed implementation optimized for the alternative hardware architecture.
This is the sort of specialized library that matters when trying to turn an AI processor into a practical development platform.
DeepEP Helps AI Chips Communicate
Another component, DeepEP, focuses on communication between computing devices.
This becomes especially important with large AI models.
Modern models are often too computationally demanding to run efficiently on a single accelerator.
Instead, workloads are distributed across dozens, hundreds or even thousands of chips.
Those processors constantly need to exchange information.
If communication is inefficient, processors can spend valuable time waiting for data instead of performing calculations.
DeepEP is designed to improve communication in distributed AI workloads, particularly systems based on mixture-of-experts architectures.
This kind of infrastructure becomes more important as Huawei builds larger systems containing many Ascend processors.
DeepSeek and Huawei Built a 128-Chip Supernode
The collaboration between DeepSeek and Huawei extends beyond software libraries.
The two companies also worked together on a supernode system based on 128 Huawei Ascend 950 processors.
According to DeepSeek, the project involved optimizing both computing and communication performance across the interconnected chips.
A supernode groups many accelerators together so that they can operate as a tightly integrated computing system.
This approach is becoming increasingly important in artificial intelligence because frontier models require enormous amounts of computing power.
Nvidia follows a similar strategy with systems that combine large numbers of accelerators through high-speed networking.
Huawei is attempting to build its own version of that vertically integrated model.
The combination of Ascend processors, supernode hardware and DeepSeek software is part of that effort.
Huawei Is Building Its Next Generation of AI Chips
The software release comes shortly after Huawei presented new generations of its Ascend processors and computing systems.
Huawei has been expanding its AI hardware roadmap as it attempts to become a larger supplier of processors for Chinese AI companies.
The company’s strategy increasingly involves not just individual chips but entire computing systems connecting many accelerators together.
DeepSeek’s new software is intended to make those platforms easier to program and more effective for demanding AI workloads.
That relationship is particularly significant because DeepSeek itself is one of China’s most prominent AI model developers.
Instead of Huawei creating software in isolation, the chipmaker is working with a company that has practical experience training and operating large language models.
US Export Controls Accelerated the Search for Alternatives
China’s efforts to build alternatives to Nvidia have accelerated under US restrictions on exports of advanced AI processors.
The controls have limited Chinese companies’ access to some of Nvidia’s most powerful chips.
That has created strong incentives for domestic companies to develop alternatives.
Huawei has emerged as one of the most prominent Chinese competitors through its Ascend family.
DeepSeek has also been increasing its use of Huawei processors, with the two companies working to improve compatibility between DeepSeek models and Ascend hardware.
But the restrictions have also highlighted an important lesson.
Hardware independence does not automatically create software independence.
A domestic chip industry still needs compilers, libraries, frameworks and developer tools.
The new open-source releases target precisely that part of the problem.
Why Open Source Matters to the Huawei Strategy
Making the DeepSeek tools open source could help the ecosystem develop faster.
Software developers can study how the tools work rather than treating them as a closed system.
Researchers can adapt them.
Companies can test them against their own workloads.
Hardware engineers can identify performance bottlenecks.
Developers can also contribute changes back to projects where the development model allows it.
This approach has already played an important role in DeepSeek’s broader strategy.
The company has released models, research and infrastructure software publicly, helping build a large developer community around its technology.
Applying the same philosophy to Huawei Ascend support could make it easier for programmers to experiment with alternatives to Nvidia hardware.
DeepSeek Open-Source Tools Do Not Replace CUDA Overnight
It would be misleading to describe the announcement as the immediate replacement of CUDA.
CUDA has a massive lead.
Millions of developers are familiar with it.
Major AI frameworks work extensively with Nvidia hardware.
Universities teach CUDA programming.
Cloud providers offer large fleets of Nvidia accelerators.
Software vendors optimize their products around Nvidia’s platform.
A six-tool open-source release does not erase that ecosystem.
The more realistic interpretation is that DeepSeek and Huawei are filling pieces that an alternative platform needs before it can become genuinely competitive.
The challenge is not simply producing equivalent tools.
Those tools also need to become reliable, well documented and widely adopted.
Developer Adoption Will Be the Bigger Test
Technology ecosystems ultimately depend on developers.
If programmers find Huawei’s tools difficult to use, adoption will remain limited regardless of theoretical performance.
If developers can move existing workloads relatively easily and achieve competitive performance, the ecosystem becomes much more attractive.
This means metrics beyond chip speed will matter.
Developers will look at documentation quality.
They will consider debugging tools.
They will need compatibility with frameworks such as PyTorch.
Companies will examine reliability and long-term support.
Researchers will compare performance across different workloads.
The ultimate battle between CUDA and alternatives will therefore take place as much in developer workflows as in semiconductor factories.
Competition Is Growing Beyond Huawei
Huawei is not the only company challenging Nvidia’s software advantage.
AMD has spent years improving its ROCm software platform for AI and high-performance computing.
Other chip companies and cloud providers are also developing accelerator software stacks.
New programming approaches such as Triton have made it easier to write high-performance kernels that can potentially operate across different hardware environments.
This means the AI industry is gradually exploring ways to make software less dependent on a single accelerator architecture.
However, Nvidia remains deeply entrenched because its hardware and software have evolved together for years.
Huawei’s advantage in China is that domestic demand for alternatives is unusually strong.
DeepSeek Could Become Important to Huawei’s AI Ecosystem
DeepSeek brings something valuable to the partnership: experience running advanced AI models.
Hardware makers can produce tools based on how they believe developers will use their chips.
AI laboratories use those systems under demanding real-world conditions.
That allows them to discover problems that may not appear in simple benchmark tests.
DeepSeek has become particularly known for designing AI architectures with computational efficiency in mind.
That makes its infrastructure software relevant to Huawei’s effort to improve Ascend performance.
If those optimizations prove useful beyond DeepSeek’s own models, other developers could benefit from the same tools.
This may help Huawei build an ecosystem influenced directly by companies creating frontier AI models.
China Is Trying to Build More of the AI Stack at Home
The DeepSeek-Huawei collaboration also illustrates a wider shift in technology competition.
The contest is no longer focused only on which country or company can produce the fastest chip.
Increasingly, it involves the entire AI stack.
That includes semiconductor design, manufacturing, networking, data centers, programming languages, software libraries, AI frameworks and the models themselves.
China has strong incentives to develop domestic capabilities across more of those layers as access to American technology becomes less predictable.
Huawei provides much of the hardware.
DeepSeek brings model development and increasingly the software infrastructure around it.
The combination illustrates how China’s AI ecosystem is becoming more vertically integrated.
A More Competitive AI Chip Market Could Benefit Developers
If Huawei and other companies build viable alternatives to CUDA, developers could ultimately gain more choice.
Competition can encourage vendors to improve software quality, reduce costs and make their platforms easier to use.
AI companies might also be able to distribute workloads across multiple kinds of accelerator rather than depending heavily on one supplier.
That could become increasingly important as global demand for AI computing continues to grow.
However, hardware diversity can also create complexity.
Developers may need to support several programming environments.
Models may behave differently across accelerators.
Optimization work may have to be repeated.
High-level tools such as TileLang attempt to reduce some of that complexity by making the underlying hardware less visible to programmers.
DeepSeek Open-Source Tools Mark a Software Push, Not Just a Chip Race
The most important part of DeepSeek’s announcement is therefore not simply that Huawei has another set of programming tools.
It is what those tools represent.
China’s effort to compete in AI computing is moving beyond the processor itself and deeper into the software ecosystem that determines whether those processors are practical to use.
DeepSeek open-source tools such as TileLang, DeepGEMM and DeepEP give Huawei’s Ascend platform more of the infrastructure required for demanding AI workloads.
CUDA remains far more mature and deeply embedded across global AI development.
But Huawei does not need to reproduce decades of Nvidia software development in a single announcement.
It needs to keep reducing the reasons developers cannot or will not move workloads onto Ascend.
The latest DeepSeek releases remove several more of those barriers.
The larger question is whether enough developers, AI laboratories and cloud providers begin using the tools for a genuine alternative software ecosystem to emerge.
If that happens, competition between Nvidia and Huawei will increasingly be decided not only by which company builds the fastest AI chip, but by which one makes those chips easiest for developers to use.







