Realtime AI News
Huawei GTS teaches its ops agent to 'watch' networks troubleshoot, nearly clears dual-firewall test
Huawei's GTS unit has trained an agent to diagnose network faults by 'watching' the network, and it nearly cleared the hard dual-firewall scenario, according to qbitai. The reported results include a 24.2% lift in task pass rate and as much as 45% lower token cost.
Network troubleshooting is turning into a proving ground for AI agents, and Huawei's Global Technical Service (GTS) unit is betting on an agent that can 'watch' a network while it diagnoses faults.
According to qbitai, the approach nearly cleared the hard dual-firewall scenario, lifting task pass rate by 24.2% and cutting token cost by as much as 45%. Better accuracy and lower cost together is exactly the signal operators need before moving agents from pilots into production.
Troubleshooting is hard because the evidence is scattered. Topology, device configuration, logs and alarms live in different systems, so engineers have to compare views, form hypotheses and verify them one by one, and a single wrong step sends the investigation down the wrong path.
The phrase 'watching the network' is the interesting part. It suggests the agent is not just replaying a text command sequence but has to make sense of interfaces and live network state, which raises the bar for multimodal perception and for the stability of tool-calling chains.
Dual-firewall setups typically involve policy interplay across more than one device, where the root cause rarely sits at a single point and has to be found by repeatedly cross-checking multiple views. That work has long depended on experienced engineers and is among the hardest things to cover with conventional automation scripts.
The 45% token-cost reduction matters just as much. Network operations tasks are frequent and long-chained, so inference cost directly determines whether a deployment can scale; a lower cost per task is what lets agents take over more routine work.
What to watch next: whether the capability generalizes beyond a specific scenario to more complex cross-domain networks, and whether Huawei GTS packages it as a product for carriers and enterprise customers.
Why it matters
If agents can reliably close the loop on high-value, high-frequency network operations, it reshapes how carriers split work between engineers and machines, and the cost drop is what makes that scale feasible.
Nearby Updates
All09/16, 09:51
Ex-Anthropic researcher alleges the lab is accelerating the AI self-improvement race
A former Anthropic researcher has publicly alleged that the company is accelerating the race toward AI self-improvement, according to a report by chosun.com. The claim lands on a lab that has built its identity around safety, and so far the report offers the allegation itself rather than evidence outsiders can check.
09/16, 11:07
Hand the memory to the CPU: Intel lays out a data-center KV cache strategy
Intel has outlined a data-center approach that offloads the growing KV cache from GPU memory to CPU-side memory and storage so GPUs can focus on generating tokens, according to a QbitAI report published on September 16. Tests cited in the report show up to about 5x faster time to first token with tiered offloading and roughly 20% to 30% less cache space through lossless hardware compression.
09/16, 11:40
Hangzhou's Liwensuo opens Lévin Harness, an agent workspace for protein design
Hangzhou-based AI protein design company Liwensuo has released Lévin Harness, an agent-centred protein design application now open to the research community with Apple-silicon Mac support. It places data, models, plugins, compute and workflows in one workspace so that literature work, tool setup, GPU jobs and result analysis can run as a repeatable loop around the models.
09/16, 11:50
Bilibili launches AI arena with 100 models competing, GPT-6 on top
Bilibili launched its AI Infinite Arena on September 16 and published a first leaderboard in which GPT-6 Astra took the top spot across evaluations from ten creators, with domestic Chinese models taking three of the top five places, according to Securities Market Weekly. The board aggregates creator-run real-world tests of more than a hundred models and will update in real time.