Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

AI infrastructure enters the self-evolving era: Tsinghua team optimizes AI with AI to build a domestic trillion-token factory

QbitAI reports that Qingmao Intelligence, a startup founded by a Tsinghua team, is applying a self-evolving "AI optimizes AI" approach to compute infrastructure, automating operator development and chip adaptation while compressing complex operator tuning from about two weeks to roughly one hour. The company has adapted dozens of domestic chips and aims to build a trillion-token-per-day factory on domestic compute for the agent era.

Published

QbitAI reported on Aug 18 that as Silicon Valley buzzes over Recursive Self-Improvement (RSI), Qingmao Intelligence (清昴智能), a startup rooted in a Tsinghua University laboratory, is applying the "AI optimizes AI" approach to compute infrastructure itself, with the goal of running a trillion-token-per-day factory on domestic chips.

RSI — AI that proposes its own questions, designs experiments, evaluates results and improves itself from feedback — is seen as the next major frontier after AI reasoning and agents. The report cites OpenAI's GPT-5.6 cutting end-to-end service cost by 20% and lifting token generation efficiency by more than 15% through RSI-related capabilities, Anthropic disclosing that over 80% of its codebase is generated by Claude, and Jeff Dean leaving Google to found Discovery Loop in the same arena.

The report highlights a mismatch in pace: upper-layer models are moving toward L4-style autonomous execution, running for hours as agents, while the software ecosystem around domestic chips remains fragmented. Operator adaptation, compilation optimization, memory scheduling and framework porting still depend on a handful of senior engineers doing manual tuning that can take weeks per complex operator — work that rapid model and chip iteration can quickly render obsolete.

Qingmao Intelligence is described as China's first startup focused on self-evolving AI infrastructure. Its technical lineage traces to the Tsinghua computer science lab of Professor Zhu Wenwu, whose team formally proposed self-driven machine learning in 2022, extending a machine's autonomy to tasks, data, models, optimization strategies and evaluation criteria. Founder Guan Chaoyu, a student of Zhu and a Tsinghua top-scholarship winner, previously led teams to championships in global AutoML and MetaDL challenges, beating Harvard, MIT and Stanford.

On the engineering side, the system splits into two capabilities: autonomous optimization, which locates bottlenecks from real running state and tries different solutions for "how to optimize this time," and autonomous evolution, which accumulates experience so future optimization across new models and chips gets faster, forming a closed loop of optimize, learn, optimize again.

The approach has been validated in real production on domestic compute: complex operator development and tuning that took a senior expert about two weeks now takes about one hour, a roughly 336x efficiency gain; full-stack enablement of a domestic chip has been compressed from about a month to one day or a week, about a 30x speedup; and some representative operators run up to 30x faster than baseline. The company says it has adapted and optimized dozens of domestic chips, covering more than half of mainstream domestic AI accelerators.

Optimizing compute is only the first step. In the agent era, some complex agents consume hundreds of times more tokens per task than traditional chat models, and daily call volumes in many business scenarios are approaching the trillion level — what decides whether a business works is the compute cost, latency and utilization behind each token. Qingmao's goal is to build a trillion-token-per-day factory on domestic chips to power that demand.

The question to watch is whether this self-evolving infrastructure paradigm spreads across the domestic compute ecosystem, and how "AI optimizing AI" rewrites the rules of the compute software stack. If it matures, optimization experience once held by a few elite engineers becomes a systematic capability that machines can execute, accumulate and replicate across models and chips.

Why it matters

If self-evolving infrastructure matures, the cost of adapting and tuning domestic chips could fall sharply, potentially redefining the competitiveness of China's domestic AI compute ecosystem.

AI Infra清昴智能国产算力
Back to realtime news

Nearby Updates

All