Realtime AI News
Zhongcheng Hualong Unveils HL200 Inference Chip and 10,000-Card Supernode Cluster
Zhongcheng Hualong held its 2026 GPU launch event in Beijing, unveiling the second-generation HL200 inference chip and a supernode computing cluster that scales to 10,240 cards. The company says single-card FP8 and FP4 compute reach 2P and 4P respectively, positioning the lineup as a full-stack domestic answer to large-scale AI inference deployment.
Zhongcheng Hualong held its 2026 GPU product launch in Beijing, officially introducing the new HL200 inference chip and a supernode computing cluster solution, saying it marks a leap for domestic AI inference compute from single-chip breakthroughs to coordinated 10,000-card-scale clusters. Bi Kaichun, executive deputy director and secretary-general of the Electronic Information Technology Committee under China's Ministry of Industry and Information Technology, and Liu Zhiguang, chairman of the China Science and Technology Consulting Association, attended and delivered remarks.
The AI industry is shifting from training-driven to inference-driven development. As AI agents scale up and long-context, multi-turn dialogue scenarios spread, token consumption is climbing fast, making efficient low-latency inference for very long contexts a core competitive moat in the compute race. Zhongcheng Hualong kept its "one chip per year" cadence and built the second-generation HL200 on its previous chip generation to target high-concurrency generative AI and agent inference workloads.
The HL200 natively supports FP4 and FP8 ultra-low-precision inference and is compatible with mainstream FP16 precision. The company reports single-card FP16/BF16 compute of 0.5P, FP8 compute of 2P, and FP4 compute of 4P, with an energy efficiency of 5.12 TFLOPS/W, which it says ranks in the domestic first tier. On the ecosystem side, the chip is compatible with PyTorch, ONNX and vLLM, and ships with an OpenAI-compatible interface and lightweight migration tooling to cut the cost of switching.
In benchmark adaptation testing against mainstream models including DeepSeek V4 Flash, DeepSeek V4 Pro and GLM 5.2, the company says HL200 leads comparable international chips across the full Prefill pipeline, with notable advantages in decode latency, and can stably support large-scale, time-sensitive commercial inference deployments.
Alongside the chip, Zhongcheng Hualong released the HL200 supernode cluster: a single cabinet interconnects 64 GPUs at high speed, a single node can stack up to 1,024 cards vertically, and horizontal scaling reaches 10,240 cards, covering everything from 8-card to 10,000-card deployments. The cluster coordinates compute, storage and networking with end-to-end SLO optimization, hitting advanced industry levels in compute density, interconnect bandwidth, deployment latency and energy efficiency.
At the event, Zhongcheng Hualong signed strategic cooperation agreements with more than ten leading companies, including China Electric Engineering under China Energy Engineering, Inspur Information, Unisplendour Smart Computing and Sinnet Cloud, to jointly build an open, controllable domestic AI compute ecosystem. Chairman Dr. Wang Jiacheng said the compute competition has moved beyond raw parameter comparisons into a comprehensive contest of inference efficiency, deployment experience and industrial value.
The company describes itself as one of the few domestic players with full-stack, independently controllable "chip + system + solution" capabilities, having built a GPU and CPU multi-core technology layout. Its core products have won bids from major operators including China Mobile and China Telecom, plus nearly 50 key contract sections across more than 20 provinces, serving critical information infrastructure in government, finance, energy and transportation.
The launch is framed as a turning point in which domestic inference compute moves from scattered single-point breakthroughs to a coordinated ecosystem spanning chips, systems, clusters and partners. The next things to watch are HL200's commercial ramp-up, its deployment progress with operators and government enterprise customers, and how far domestic inference chips can close the gap in energy efficiency and ecosystem fit.
Why it matters
The HL200 and supernode cluster give domestic AI inference a path from single-chip performance to 10,000-card-scale supply, intensifying competition in energy efficiency, ecosystem fit and scaled deployment, while offering Chinese enterprises a new independently controllable option for commercializing trillion-parameter models.
Nearby Updates
All08/24, 14:51
Data Robotics unveils Hairui Cloud Brain 6.0 and Taishi OS 1.0, touting an “Android moment” for robotics
On August 24, Data Robotics released Hairui Cloud Brain 6.0 and the Taishi OS 1.0 operating system. The company, which raised a 3 billion yuan seed round — the largest for a Chinese tech startup — framed the launch as an “Android moment” for the robotics industry.
08/24, 15:51
Yonyou's H1 2026 Report Signals Enterprise AI Moving into Large-Scale Deployment
Yonyou Network published its 2026 semi-annual report, stating that enterprise AI has entered the stage of large-scale deployment. The signal from one of China's leading enterprise software vendors suggests corporate AI is shifting from pilots to systemic rollout across core business functions.
08/24, 15:54
Nvidia in talks to invest in Perplexity at a $30B+ valuation
Nvidia is in talks to invest in AI search startup Perplexity as part of an equity round valuing the company at more than $30 billion, The Information reported on Sunday. The valuation would be up more than 50% from a year ago, as Perplexity's annualized revenue has climbed past $750 million.
08/24, 13:35
Micron CEO Says He Sees 'No End' to AI Memory Demand
Micron's CEO said publicly that AI-driven demand for memory chips has no end in sight, signaling strong confidence in the sector's outlook. The comments land as AI data center buildouts become the strongest demand driver for memory.