Realtime AI News
Chinese supercomputer runs full DeepSeek-V3/R1 models on pure CPUs, matching an 80-GPU cluster
Reports say the LingSheng supercomputer, ranked the world's top domestic Chinese supercomputer, has completed distributed inference of MoE models on a pure CPU architecture, running the full DeepSeek-V3/R1-671B model on just 16 compute nodes. At a batch size of 2048, its output throughput is said to be comparable to a cluster of 80 mainstream GPUs.
The traditional assumption that running large language models requires GPUs is being challenged. According to reports, the LingSheng supercomputer — ranked the world's top domestic Chinese supercomputer — has completed distributed inference deployment of MoE models on a pure CPU architecture, running the full DeepSeek-V3/R1-671B model smoothly on just 16 compute nodes.
At a batch size of 2048, the supercomputer's output throughput is said to be comparable to a cluster of 80 mainstream GPUs. If accurate, it would mark the first time a pure domestic CPU architecture has competed head-on with GPU clusters in large-model inference.
The LingSheng supercomputer's most distinctive feature is its pure CPU architecture. Instead of the traditional approach of stacking GPUs or dedicated accelerator cards, its self-developed LX2 CPU integrates matrix acceleration units directly on-chip, supporting multi-precision computing so each CPU can handle both supercomputing science workloads and AI inference.
This on-chip fusion design eliminates the data movement overhead between CPU and accelerator cards. But compute is only the first step — the real bottleneck in large-model inference is memory bandwidth, since model parameters must be shuttled constantly between compute units and memory. To address this, the LX2 CPU integrates China's first domestic high-bandwidth memory (HBM), with total bandwidth of 4TB/s, a 10x improvement over traditional CPU memory bandwidth.
Paired with the self-developed LingQi high-speed network supporting microsecond-level latency and 100,000-node networking, the design clears the three major hurdles of compute, bandwidth, and communication.
The progress matters for China's domestic compute ecosystem: the pure CPU route offers an alternative domestic path for large-model inference that does not depend on GPUs, turning the idea that model compute no longer depends on graphics cards from a slogan into a verifiable engineering practice.
That said, the figures currently come mainly from media reports, and no official benchmark from the LingSheng supercomputer team has been published yet. The key things to watch are whether the pure-CPU inference approach can be reproduced in larger commercial inference scenarios, and the production progress of the LX2 CPU and domestic HBM.
Why it matters
If the reported figures hold, LingSheng's pure-CPU deployment of the full DeepSeek-V3/R1 models offers an alternative path for domestic large-model inference without GPUs. Watch for official benchmarks and the production rollout of the LX2 CPU and domestic HBM.
Nearby Updates
All08/08, 17:40
Apple Intelligence officially supports Alibaba's Qwen models in deep tech collaboration
Apple Intelligence has officially added support for Alibaba's Qwen large language models, with the two companies forming a deep technical collaboration, according to a SmartHey report. The move marks a key step in localizing Apple's AI services in China.
08/08, 18:24
Firebird launches CIS region's largest AI factory in Armenia, powered by NVIDIA and Dell
Firebird, an emerging AI cloud provider, has launched the CIS region's largest AI factory in Armenia, creating a new AI computing hub powered by NVIDIA accelerated computing and Dell Technologies high-performance infrastructure. Armenian Prime Minister Nikol Pashinyan and Deputy Prime Minister Zhaslan Madiyev attended the launch, underscoring the project's regional significance.
08/08, 16:37
Gemini App reaches 950 million monthly users, Google says
Google says its Gemini app has reached 950 million monthly active users. The milestone makes Gemini one of the largest consumer AI assistant apps in the world and a strong signal for mainstream adoption of generative AI.
08/08, 13:59
Alibaba unveils Qwen3.8-Max AI model to drive cloud growth
Alibaba has unveiled a new large language model, Qwen3.8-Max, in a move widely seen as an effort to drive cloud growth with stronger AI capabilities. Public details remain thin, with parameters, benchmarks and pricing still awaiting official disclosure.