Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

Zhongcheng Hualong Unveils HL200 Inference Chip and 10,000-Card Supernode Cluster

Zhongcheng Hualong held its 2026 GPU launch event in Beijing, unveiling the second-generation HL200 inference chip and a supernode computing cluster that scales to 10,240 cards. The company says single-card FP8 and FP4 compute reach 2P and 4P respectively, positioning the lineup as a full-stack domestic answer to large-scale AI inference deployment.

Published

Zhongcheng Hualong held its 2026 GPU product launch in Beijing, officially introducing the new HL200 inference chip and a supernode computing cluster solution, saying it marks a leap for domestic AI inference compute from single-chip breakthroughs to coordinated 10,000-card-scale clusters. Bi Kaichun, executive deputy director and secretary-general of the Electronic Information Technology Committee under China's Ministry of Industry and Information Technology, and Liu Zhiguang, chairman of the China Science and Technology Consulting Association, attended and delivered remarks.

The AI industry is shifting from training-driven to inference-driven development. As AI agents scale up and long-context, multi-turn dialogue scenarios spread, token consumption is climbing fast, making efficient low-latency inference for very long contexts a core competitive moat in the compute race. Zhongcheng Hualong kept its "one chip per year" cadence and built the second-generation HL200 on its previous chip generation to target high-concurrency generative AI and agent inference workloads.

The HL200 natively supports FP4 and FP8 ultra-low-precision inference and is compatible with mainstream FP16 precision. The company reports single-card FP16/BF16 compute of 0.5P, FP8 compute of 2P, and FP4 compute of 4P, with an energy efficiency of 5.12 TFLOPS/W, which it says ranks in the domestic first tier. On the ecosystem side, the chip is compatible with PyTorch, ONNX and vLLM, and ships with an OpenAI-compatible interface and lightweight migration tooling to cut the cost of switching.

In benchmark adaptation testing against mainstream models including DeepSeek V4 Flash, DeepSeek V4 Pro and GLM 5.2, the company says HL200 leads comparable international chips across the full Prefill pipeline, with notable advantages in decode latency, and can stably support large-scale, time-sensitive commercial inference deployments.

Alongside the chip, Zhongcheng Hualong released the HL200 supernode cluster: a single cabinet interconnects 64 GPUs at high speed, a single node can stack up to 1,024 cards vertically, and horizontal scaling reaches 10,240 cards, covering everything from 8-card to 10,000-card deployments. The cluster coordinates compute, storage and networking with end-to-end SLO optimization, hitting advanced industry levels in compute density, interconnect bandwidth, deployment latency and energy efficiency.

At the event, Zhongcheng Hualong signed strategic cooperation agreements with more than ten leading companies, including China Electric Engineering under China Energy Engineering, Inspur Information, Unisplendour Smart Computing and Sinnet Cloud, to jointly build an open, controllable domestic AI compute ecosystem. Chairman Dr. Wang Jiacheng said the compute competition has moved beyond raw parameter comparisons into a comprehensive contest of inference efficiency, deployment experience and industrial value.

The company describes itself as one of the few domestic players with full-stack, independently controllable "chip + system + solution" capabilities, having built a GPU and CPU multi-core technology layout. Its core products have won bids from major operators including China Mobile and China Telecom, plus nearly 50 key contract sections across more than 20 provinces, serving critical information infrastructure in government, finance, energy and transportation.

The launch is framed as a turning point in which domestic inference compute moves from scattered single-point breakthroughs to a coordinated ecosystem spanning chips, systems, clusters and partners. The next things to watch are HL200's commercial ramp-up, its deployment progress with operators and government enterprise customers, and how far domestic inference chips can close the gap in energy efficiency and ecosystem fit.

Why it matters

The HL200 and supernode cluster give domestic AI inference a path from single-chip performance to 10,000-card-scale supply, intensifying competition in energy efficiency, ecosystem fit and scaled deployment, while offering Chinese enterprises a new independently controllable option for commercializing trillion-parameter models.

中诚华隆推理芯片算力基础设施
Back to realtime news

Nearby Updates

All