Realtime AI News
DeepSeek open-sources Ascend infrastructure: TileLang, DeepGEMM and DeepEP
On September 30, DeepSeek open-sourced infrastructure components for Huawei's Ascend platform, including the TileLang high-level compiler toolchain, high-performance compute libraries and a distributed communication library that mirror its earlier GPU releases. The batch includes DeepGEMM, FlashMLA, TileKernel and DeepSelect operator libraries plus DeepEP, while Huawei opened their joint deployment work in the CANN community.
DeepSeek officially open-sourced a set of infrastructure components for Huawei's Ascend compute platform on September 30, covering the TileLang high-level language compiler toolchain along with high-performance compute libraries and a distributed communication library, according to Chinese tech outlet QbitAI. The release mirrors components DeepSeek had already open-sourced for GPU platforms, and the report calls it a milestone for China's AI software ecosystem because it gives developers a genuine alternative compute stack.
The release includes high-performance Ascend operator libraries such as DeepGEMM, FlashMLA, TileKernel and DeepSelect, plus the DeepEP distributed communication library. In effect, the performance stack DeepSeek proved out on GPUs is being systematically ported to Ascend.
On the toolchain side, Ascend offers a stable and open Ascend C API so developers can tune memory access paths and compute pipelines when writing high-performance operators. Building on that API and the PTO ISA instruction layer, Ascend also supports DeepSeek's TileLang programming work. The ecosystem is meant to serve both paradigms: hand-tuned optimization by experienced engineers, and compilation from high-level languages such as TileLang.
For hardware and networking, Huawei provided DeepSeek with jointly defined Ascend supernode designs, SuperPoD Flex and the UBL128 interconnect, capable of a 128-card, 3.2Tbps single-layer scale-up network and a 256K-card two-layer scale-out network, aimed at ultra-low-latency inference and large-scale training of frontier base models. Ascend also supplies ASC-COMM, a high-performance library for custom communication programming.
Communication is the other centerpiece. DeepEP, developed by the DeepSeek team, covers communication operators across EP, CP, PP and FSDP modes to support scaling larger models onto larger clusters, and measured interconnect bandwidth reaches 375 GB/s for dispatch and 347 GB/s for combine — close to the hardware limit.
To let developers deploy DeepSeek models on Ascend 950 and supernode clusters, Huawei open-sourced the joint work in the CANN community, spanning large-EP low-latency inference deployment, single-card and single-machine deployment, large-scale training, long-context KV-cache pooling and agentic RL. In large-scale inference, using an EP32 deployment strategy in offline mode, DeepSeek-V4.1-Flash reaches 2,469 output tokens per second per card at a TPOT of 5ms, and 5,102 tokens per second per card at a TPOT of 10ms without a serving framework.
The report notes those benchmark figures were collected in offline inference mode and exclude serving scheduling and framework load-balancing effects, with a context length of 128K and a Dspark speculative acceptance rate of 0.85. The article was supplied by Huawei and republished by QbitAI with permission, so it reflects a vendor's framing — which itself signals how aggressively chip makers and model teams are courting developer confidence in domestic compute software stacks.
The real signal is the "chip-model co-design" approach. When a leading model team open-sources its training, inference, communication and compilation components directly onto a domestic compute platform, the cost of migrating and tuning models there falls sharply. The next things to watch are how fast these components iterate in the CANN community, and whether more model teams follow by syncing their core operator libraries to Ascend.
Why it matters
The release moves DeepSeek's GPU-proven operator and communication stack onto Ascend, giving China's domestic compute platforms a usable software foundation. The deeper shift is tighter binding between leading model teams and chip makers, which could accelerate migration away from the CUDA ecosystem.
Nearby Updates
All09/30, 11:00
ByteDance's Doubao reportedly prepping a personal AI agent codenamed Spell
TechNode reports that ByteDance's Doubao is preparing a personal AI agent codenamed Spell, with the project said to have been in testing since April. The report is framed as unconfirmed, and ByteDance has not officially announced or acknowledged the product.
09/30, 09:42
Anthropic Warns of 'Existential Risks to Humanity' in Its IPO Filing
The Times of India reports that Anthropic warned in its IPO filing that AI may pose existential risks to humanity. A frontier lab putting that language into a capital markets document makes risk disclosure part of its public listing story.
09/30, 09:04
Okta's agent gateway brings runtime control to AI agents
Okta has introduced an Agent Gateway that polices what AI agents do at runtime, according to SiliconANGLE. The move targets a practical enterprise fear: agents hold real credentials and call internal systems automatically, so controlling their actions requires more than login-time identity checks.
09/30, 09:01
Democrats press for an AI testing regime
Democratic lawmakers in the United States are pressing for a formal AI testing regime, according to The Washington Post's AI and tech brief. The demand raises concrete questions about who would evaluate models before deployment, what the tests would cover, and whether results would be published.