Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

Massive 2.8-trillion-parameter model crashes GPU capacity within 48 hours, halts paid service

A new large language model with 2.8 trillion parameters — benchmarked against Claude — went viral immediately upon launch, exhausting GPU compute resources within just 48 hours. The operator was forced to suspend paid access as demand overwhelmed available infrastructure.

Published
2.8万亿参数大模型上线48小时即遭疯抢,GPU算力告急暂停付费
Image source: nvidia.com

A newly launched large language model boasting 2.8 trillion parameters has caused a frenzy among users, exhausting GPU compute capacity within just 48 hours of going live, according to reports from Chinese media. The model, described as directly competitive with Anthropic's Claude series, was so popular that its operator had to temporarily suspend paid service.

The 2.8-trillion-parameter scale far exceeds most mainstream models, enabling more complex reasoning and broader knowledge representation. However, this scale also demands enormous inference compute, and the operator's GPU cluster was quickly overwhelmed by the surge in user traffic.

The incident vividly illustrates a growing pain point in the AI industry: the rapid advancement of frontier model capabilities is outstripping the expansion rate of inference infrastructure. Even when the model delivers, the compute runs out.

Notably, the model is positioned against Claude rather than GPT or Gemini, suggesting a deliberate differentiation strategy. Claude's reputation for safety and thoughtfulness may have informed the developer's choice of benchmark target.

This episode serves as a stark reminder that despite accelerating model innovation, compute bottlenecks remain one of the most critical constraints on AI deployment at scale. The industry must find ways to balance model scale with infrastructure economics.

Why it matters

The GPU capacity crisis triggered by this massive model launch highlights the widening gap between AI model capability and the infrastructure needed to serve it at scale.

Large Language ModelAI InfrastructureGPU Computing
Back to realtime news

Nearby Updates

All