Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

QuJing Technology and Moore Threads sign strategic partnership to scale domestic AI token production

QuJing Technology and Moore Threads signed a strategic cooperation agreement on September 3, combining QuJing's domestic PD heterogeneous technology with Moore Threads' MTT S5000 cards and MUSA software to build domestic high-quality AI token production infrastructure. The joint solution is already in production carrying real token traffic from leading model vendors, claiming a cost-performance advantage over international advanced compute under the same service standards.

Published

On September 3, QuJing Technology signed a strategic cooperation agreement with Moore Threads to deeply integrate QuJing's self-developed domestic PD heterogeneous technology with Moore Threads' MTT S5000 AI computing cards and MUSA software platform. Under the deal, the two companies will advance domestic high-quality AI token factory construction on QuJing's ATaaS platform and jointly develop and promote joint solutions built around Token Pod, or token supernodes. QbitAI reports the solution is already in production, carrying real token traffic from leading model vendors' official services.

At the same production service standards, the setup delivers production-grade performance from domestic silicon plus a per-token cost advantage over international advanced compute. Infrastructure competition is shifting from single-card performance to system-level token production efficiency, the report argues, as large-model applications move into scaled commercial deployment.

The approach centers on a Prefill-Decode split: Moore Threads' MTT S5000 handles input computation and KV cache generation during Prefill, working alongside high-bandwidth GPUs that generate tokens during Decode. With dense AI compute, native FP8 support and the mature MUSA software stack, the S5000 lets Prefill be decoupled from the traditional inference pipeline into a resource pool that can be supplied, billed and optimized independently — reducing the burden on expensive resources and avoiding whole-cluster provisioning for single-phase peak demand.

After multiple rounds of performance iteration and joint optimization, the MTT S5000 has moved past model-compatibility and validation work and formally entered the high-quality AI token production pipeline, according to the report. QuJing attributes the result to its "fewer models, deeper optimization" strategy: focusing engineering on models with clear scaled demand, then feeding model configurations, traffic profiles and serving strategies back into the ATaaS platform as reusable production know-how.

At the signing ceremony, Liu Xianhe, general manager of QuJing's token business unit, and Fei Jingran, deputy general manager of Moore Threads' strategic sales division, signed on behalf of their companies, joined by Tsinghua professor and QuJing chief scientist Wu Yongwei, QuJing founder and CEO Ai Zhiyuan, and Moore Threads founder, chairman and CEO Zhang Jianzhong.

The next phase is standardizing delivery around Token Pod. The pod integrates shared KV cache, traffic-aware configuration, elastic scaling and fault isolation, and can adapt to both domestic and international mainstream hardware combinations to form configurable, expandable token production units. QuJing says it already maintains daily high-quality AI token production at the hundred-billion-token scale across multiple projects, giving the two sides operating experience for building larger domestic token factories.

The broader signal is that domestic chips are now competing on system-level token production economics rather than single-card specifications alone, though cost-effectiveness must ultimately be proven in real business. One caveat: the article was provided by QuJing Technology and republished by QbitAI, so it reflects the company's own framing; watch whether Token Pod deployments replicate across internet companies, frontier foundation-model firms and carriers.

Why it matters

The deal pushes the domestic-chip narrative from raw single-card specs to system-level token production economics, putting Moore Threads' MTT S5000 into real production traffic. The open question is whether the Token Pod model can be standardized and replicated across carriers, foundation-model firms and internet companies.

摩尔线程趋境科技国产算力AI基础设施
Back to realtime news

Nearby Updates

All

09/04, 17:19

Astribot releases SmoothRL, an asynchronous online RL framework for robots that can't wait for models

Astribot's foundation-model team has released SmoothRL, an online reinforcement learning framework designed for asynchronous execution, in which robots keep moving while the model computes the next action chunk in the background. In real-robot tests on the S1, throwing success rose from 39% to 94%, pen capping from 8% to 83%, and box opening reached 90%.

09/04, 15:55

Zhipu opens official Tmall flagship store to sell GLM Coding Plan subscriptions

Zhipu has opened an official flagship store on Alibaba's Tmall, listing GLM Coding Plan subscriptions in Lite, Pro, Max and team tiers with monthly, quarterly and yearly billing at prices matching its official site. The move shows large-model capabilities being packaged as standardized digital goods after API and open-platform revenue grew to 86.5 percent of Zhipu's first-half revenue.

09/04, 13:57

Meta's Hatch AI Agent Exposes Security Flaws During Testing

Security testing of Meta's Hatch AI agent has exposed flaws, according to a report from The Chosun Ilbo. The episode underscores how agent safety validation is struggling to keep pace with the rapid adoption of AI agents.

09/04, 13:48

Qwen Office tops 30 million users in its first month, enterprise accounts over half

Alibaba announced on September 4 that Qwen Office, its enterprise-grade general-purpose Agent product, surpassed 30 million total users within its first month after launching on August 3, with enterprise accounts making up more than half of the user base. The milestone came alongside roughly 120 version updates, an open-sourced context infrastructure called MyContext, and a dedicated office-tuned model built on Qwen3.8-Flash.