Realtime AI News
AI infrastructure enters the self-evolving era: Tsinghua team optimizes AI with AI to build a domestic trillion-token factory
QbitAI reports that Qingmao Intelligence, a startup founded by a Tsinghua team, is applying a self-evolving "AI optimizes AI" approach to compute infrastructure, automating operator development and chip adaptation while compressing complex operator tuning from about two weeks to roughly one hour. The company has adapted dozens of domestic chips and aims to build a trillion-token-per-day factory on domestic compute for the agent era.
QbitAI reported on Aug 18 that as Silicon Valley buzzes over Recursive Self-Improvement (RSI), Qingmao Intelligence (清昴智能), a startup rooted in a Tsinghua University laboratory, is applying the "AI optimizes AI" approach to compute infrastructure itself, with the goal of running a trillion-token-per-day factory on domestic chips.
RSI — AI that proposes its own questions, designs experiments, evaluates results and improves itself from feedback — is seen as the next major frontier after AI reasoning and agents. The report cites OpenAI's GPT-5.6 cutting end-to-end service cost by 20% and lifting token generation efficiency by more than 15% through RSI-related capabilities, Anthropic disclosing that over 80% of its codebase is generated by Claude, and Jeff Dean leaving Google to found Discovery Loop in the same arena.
The report highlights a mismatch in pace: upper-layer models are moving toward L4-style autonomous execution, running for hours as agents, while the software ecosystem around domestic chips remains fragmented. Operator adaptation, compilation optimization, memory scheduling and framework porting still depend on a handful of senior engineers doing manual tuning that can take weeks per complex operator — work that rapid model and chip iteration can quickly render obsolete.
Qingmao Intelligence is described as China's first startup focused on self-evolving AI infrastructure. Its technical lineage traces to the Tsinghua computer science lab of Professor Zhu Wenwu, whose team formally proposed self-driven machine learning in 2022, extending a machine's autonomy to tasks, data, models, optimization strategies and evaluation criteria. Founder Guan Chaoyu, a student of Zhu and a Tsinghua top-scholarship winner, previously led teams to championships in global AutoML and MetaDL challenges, beating Harvard, MIT and Stanford.
On the engineering side, the system splits into two capabilities: autonomous optimization, which locates bottlenecks from real running state and tries different solutions for "how to optimize this time," and autonomous evolution, which accumulates experience so future optimization across new models and chips gets faster, forming a closed loop of optimize, learn, optimize again.
The approach has been validated in real production on domestic compute: complex operator development and tuning that took a senior expert about two weeks now takes about one hour, a roughly 336x efficiency gain; full-stack enablement of a domestic chip has been compressed from about a month to one day or a week, about a 30x speedup; and some representative operators run up to 30x faster than baseline. The company says it has adapted and optimized dozens of domestic chips, covering more than half of mainstream domestic AI accelerators.
Optimizing compute is only the first step. In the agent era, some complex agents consume hundreds of times more tokens per task than traditional chat models, and daily call volumes in many business scenarios are approaching the trillion level — what decides whether a business works is the compute cost, latency and utilization behind each token. Qingmao's goal is to build a trillion-token-per-day factory on domestic chips to power that demand.
The question to watch is whether this self-evolving infrastructure paradigm spreads across the domestic compute ecosystem, and how "AI optimizing AI" rewrites the rules of the compute software stack. If it matures, optimization experience once held by a few elite engineers becomes a systematic capability that machines can execute, accumulate and replicate across models and chips.
Why it matters
If self-evolving infrastructure matures, the cost of adapting and tuning domestic chips could fall sharply, potentially redefining the competitiveness of China's domestic AI compute ecosystem.
Nearby Updates
All08/18, 11:04
Musk moves into AI coding, Microsoft shares tumble
Elon Musk has moved into AI coding, and Microsoft shares tumbled on the news, according to reports carried by Chinese financial media. The market reaction signals that AI coding — one of the most contested battlegrounds in commercial AI — is being repriced as a heavyweight entrant arrives.
08/18, 11:01
WeCom fully upgrades CLI and MCP, mainstream AI agents can connect directly
WeCom, the enterprise messaging platform, has fully upgraded its CLI and MCP capabilities, letting mainstream AI agents connect directly, according to Sina Finance. The move opens a standardized gateway for AI agents into a high-frequency workplace channel.
08/18, 10:49
Zhipu releases new-generation foundation model, tops open-source benchmarks
Zhipu has released a new generation of its foundation model, claiming top open-source results across multiple benchmarks, according to a Sohu report. The Chinese AI lab is extending its open-weights strategy, though the model's official name and full evaluation details have yet to be disclosed.
08/18, 10:33
Tsinghua-linked AI agent startup Dudao Tech raises nearly ¥100M Series A+ led by SAIF
Dudao Tech, a Tsinghua-affiliated AI agent startup, has raised a nearly 100 million yuan Series A+ round led by SAIF. The company is building an "agent compiler" aimed at turning business requirements into executable AI agents for enterprise use.