Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

Chinese supercomputer runs full DeepSeek-V3/R1 models on pure CPUs, matching an 80-GPU cluster

Reports say the LingSheng supercomputer, ranked the world's top domestic Chinese supercomputer, has completed distributed inference of MoE models on a pure CPU architecture, running the full DeepSeek-V3/R1-671B model on just 16 compute nodes. At a batch size of 2048, its output throughput is said to be comparable to a cluster of 80 mainstream GPUs.

Published

The traditional assumption that running large language models requires GPUs is being challenged. According to reports, the LingSheng supercomputer — ranked the world's top domestic Chinese supercomputer — has completed distributed inference deployment of MoE models on a pure CPU architecture, running the full DeepSeek-V3/R1-671B model smoothly on just 16 compute nodes.

At a batch size of 2048, the supercomputer's output throughput is said to be comparable to a cluster of 80 mainstream GPUs. If accurate, it would mark the first time a pure domestic CPU architecture has competed head-on with GPU clusters in large-model inference.

The LingSheng supercomputer's most distinctive feature is its pure CPU architecture. Instead of the traditional approach of stacking GPUs or dedicated accelerator cards, its self-developed LX2 CPU integrates matrix acceleration units directly on-chip, supporting multi-precision computing so each CPU can handle both supercomputing science workloads and AI inference.

This on-chip fusion design eliminates the data movement overhead between CPU and accelerator cards. But compute is only the first step — the real bottleneck in large-model inference is memory bandwidth, since model parameters must be shuttled constantly between compute units and memory. To address this, the LX2 CPU integrates China's first domestic high-bandwidth memory (HBM), with total bandwidth of 4TB/s, a 10x improvement over traditional CPU memory bandwidth.

Paired with the self-developed LingQi high-speed network supporting microsecond-level latency and 100,000-node networking, the design clears the three major hurdles of compute, bandwidth, and communication.

The progress matters for China's domestic compute ecosystem: the pure CPU route offers an alternative domestic path for large-model inference that does not depend on GPUs, turning the idea that model compute no longer depends on graphics cards from a slogan into a verifiable engineering practice.

That said, the figures currently come mainly from media reports, and no official benchmark from the LingSheng supercomputer team has been published yet. The key things to watch are whether the pure-CPU inference approach can be reproduced in larger commercial inference scenarios, and the production progress of the LX2 CPU and domestic HBM.

Why it matters

If the reported figures hold, LingSheng's pure-CPU deployment of the full DeepSeek-V3/R1 models offers an alternative path for domestic large-model inference without GPUs. Watch for official benchmarks and the production rollout of the LX2 CPU and domestic HBM.

DeepSeekSupercomputerCPU Inference
Back to realtime news

Nearby Updates

All