Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

openJiuwen debuts X-Router self-evolving model routing with Ascend affinity, cutting agent token consumption by over 50%

The openJiuwen team has released X-Router, a self-evolving model routing technology that adds a request-level decision layer between agents and their models, sending simple tasks to lightweight models and escalating only hard ones. In official tests on PinchBench and Terminal-Bench, enabling self-evolution cut model-call costs by as much as 51.4% while matching the all-cloud strong-model baseline.

Published

On October 2 the open-source AI agent platform openJiuwen released X-Router, a self-evolving model routing technology that inserts a request-level decision layer between the agent and its pool of candidate models. openJiuwen is built jointly by teams including Huawei's 2012 Labs, Huawei Cloud, and its device, computing and computing-power units together with universities and outside developers, and X-Router is the first systematic public disclosure of that routing capability.

As the report frames it, most AI applications still rely on one model for everything: a request to translate a sentence and a request that requires writing code and debugging it three times both land on the same hundred-billion-parameter model. Once an agent holds local models, cloud models and services from several vendors at the same time, a hard-coded routing table cannot keep up with model iterations and cannot balance cost against quality.

openJiuwen splits routing into three layers. The first builds capability profiles for each model, drawing them offline and refreshing them online across task types, difficulty levels, latency, power and cost, with profile updates landing in under a minute when a model version or its behaviour changes. The second handles the decision itself, taking in not just the complexity of the current request but the whole multi-turn trajectory and the state of the system right now. The third makes the policy evolve as real feedback accumulates.

The decision layer folds in signals that pure algorithms usually miss. It reads KV-cache affinity to judge which model already holds a warm cache for this context, checks real-time load so a request is not pushed onto a busier model, and writes recently timed-out or rate-limited models into state as exclusions for the next round. The algorithm itself is a pure function: the same request and the same state snapshot must produce the same decision, and all cross-request memory lives in a separate state layer, so losing state only downgrades to "cold routing" instead of failing the request.

Above model selection, the router can orchestrate multi-model collaboration. When one model is not enough for a trustworthy answer, the routing layer decides whether to launch several models this round and which ones, and WorkSwarm's MoA mechanism handles semantic de-duplication, quality filtering, conflict resolution and compression. The team describes the granularity shifting from "choosing a model" to "orchestrating the whole execution strategy", placing models, tools and sub-agents inside one scheduling framework.

The measurements come from x-router plugged into WorkSwarm and evaluated on the full 147-task PinchBench set, spanning 11 categories including log analysis, data analysis, coding and research, with a locally deployed Qwen3-0.6B classifier running in-process. Sending everything to the cloud model Kimi-K2-Thinking scored 71.36% at a cost of $11.11; x-router's static routing scored 66.3% at $8.25, cutting cost by 25.7%; switching on the Bandit self-evolution mode pushed cost down to $6.16 while the score recovered to 70.70% - roughly 44.6% cheaper than the baseline.

A second test on Terminal-Bench/LLMRouterBench used a three-tier pool of Opus4.8, Qwen3.5-122B and Qwen3.5-35B. In cost-focused mode it cut cost by 51.4% while reaching 95.2% of Opus's success rate; in quality-focused mode it cut cost by 15.7% while reaching 98.6%. Both benchmarks point the same way: near-top-tier quality at roughly half the cost, with how much quality to trade for how much saving left as a user-configurable choice.

For agents, routing is turning from a nice-to-have optimization into a precondition for scale. Unlike a chatbot, an agent works continuously across many turns and tools, so token consumption multiplies rather than growing linearly; openJiuwen says the kernel aims for generality rather than a single-point optimum, with one framework configurable into edge, cloud, on-premise and hybrid deployments. The core is written in Rust with a Python layer via PyO3, and the code is open-sourced in the openJiuwen/model-router repository on gitcode.

Why it matters

Model routing turns the cost-versus-quality trade-off into a configurable, evolvable and observable system capability, and that is the difference between an agent that only works in a demo and one that survives real production traffic. As the model ecosystem keeps fragmenting, choosing the right model is becoming infrastructure in its own right.

openJiuwenAgentModel RoutingAscend
Back to realtime news

Nearby Updates

All

10/02, 15:35

After Meta, Manus Regains Independence and Debuts Manus 2.0 and Cue Agent

Manus has regained its independence after its time under Meta and used the moment to unveil Manus 2.0 and a new AI agent called Cue, according to digitimes. The twin release signals that the company intends to keep shipping in the crowded general-agent market on its own terms.

10/02, 15:01

Pi coding agent reverses course and adds MCP support with 1.0 release

The Pi coding agent hit its 1.0 milestone on Thursday and added Model Context Protocol (MCP) support to its core, reversing creator Mario Zechner's earlier position that MCP was unnecessary. Earendil, the company that acquired Pi in April 2026, said MCP has improved and that its own MCP changes make it easier to integrate other capabilities, including the Jev decision model.

10/02, 14:46

arXiv Caps Submissions at Two Per Month, With Rejections Still Counting

Starting October 1, 2026, arXiv limits each submitter to two paper submissions per calendar month across all categories, and rejected papers still consume the monthly quota. QbitAI reports that the change is meant to ease the review strain caused by a surge in submissions and low-value papers in the AI era.

10/02, 14:43

Yokogawa joins Singapore AI data-centre testbed consortium

Yokogawa has joined a Singapore-based AI data-centre testbed consortium, according to IT Brief Asia. The move highlights how AI compute build-out is converging with data-centre infrastructure, and reflects Singapore's continued push to anchor the region's AI ecosystem.