DeepSeek, China’s answer to OpenAI, has teamed up with Huawei to develop programming tools for its Ascend chips.
Reuters cites a post on DeepSeek’s WeChat account confirming the partnership, with tools used to develop on Huawei's Ascend line, including compilers and related libraries, to be open-sourced.
The programming infrastructure is specially designed for Huawei hardware, including the Ascend 950 chip – the successor to its 910C AI chip – with the pair believed to have worked on a “supernode” iteration that sees 128 Ascend 950s connected.
DeepSeek’s decision to create optimized programming tools for developing on Huawei hardware follows prior claims that users in China criticized its CUDA alternative, dubbed CANN, with reports that some of its own staff called it “difficult and unstable to use.”
In addition to optimizing its own compilers for use on Huawei hardware, DeepSeek has thrown its weight behind TileLang, an open-source composable tiled programming model developed by Microsoft researchers that works similarly to Nvidia’s CUDA Tile. The Chinese lab’s post on WeChat suggests TileLang provides “a simpler programming model” than CUDA.
“To build a new generation of independent, self-controlled GPU software ecosystems, the first priority is establishing a high-level language that is universal, easy to program, and still capable of reaching the hardware's full performance potential,” DeepSeek said.
With export restrictions hampering its ability to access high-end Western hardware, DeepSeek’s whole strategy has been software optimization, getting the best out of the constrained chips available. Its entire dramatic entrance into the AI zeitgeist was built on using older-generation Nvidia hardware to train models on par in terms of performance with frontier systems at the time.
DeepSeek has continued to iterate, culminating in Engram. Described as a "conditional memory" concept, it bypasses GPU memory constraints by offloading tactical knowledge (simple information lookups) to the CPU to relieve an AI model’s core computational network, thereby allowing it to focus on more complex reasoning tasks.
Away from optimizations, the Chinese lab has sought to entice Chinese developers – and those further afield – by dramatically slashing API costs by up to 90%, at a time when token costs are spiraling.
Comments