Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

Alibaba's Qwen3.8 tops Artificial Analysis leaderboard for Agentic capability

Alibaba's Qwen3.8 has taken the top spot globally for Agentic capability on the Artificial Analysis leaderboard, according to a report from Chinese tech media QbitAI. The result places the open-source model ahead of leading domestic and international rivals in agent-related performance, marking another milestone for the Qwen family.

Published

Alibaba's Qwen3.8 has ranked first globally in Agentic capability on the Artificial Analysis leaderboard, according to a report from Chinese tech media QbitAI. The result places the open-source model ahead of leading domestic and international large language models in agent-related performance.

Artificial Analysis is a widely used third-party platform for cross-comparing large models, covering dimensions such as language, reasoning, coding, math and agentic capability. The Agentic score evaluates how well a model handles tool calls, multi-step planning and autonomous execution in real-world tasks, a key measure of whether a model can actually get work done.

As the newest open-source model from Alibaba's Tongyi Lab, Qwen3.8 builds on the influence of the Qwen family in the open-source ecosystem. Previous Qwen models have topped several public leaderboards, and this Agentic result strengthens their standing in agent application scenarios.

Agentic capability is becoming the next major battleground in the large model race. As AI agents move into production, a model's ability to reliably call tools and complete multi-step tasks directly determines the quality of downstream applications. Topping this dimension suggests Chinese open-source models can now compete head-on with the best international models at the agent infrastructure level.

For developers, the ranking offers a reference point for model selection. Models with high Agentic scores tend to mean lower debugging costs and more consistent task completion when building agent applications, and Qwen3.8's open-source nature lowers the barrier to adoption.

It is worth noting that leaderboard results reflect performance on specific evaluation dimensions, and real-world experience can vary across scenarios. Whether Qwen3.8 maintains its lead in further evaluations and production use remains to be seen.

The next points to watch are whether Alibaba ships a stronger commercial version built on Qwen3.8, and whether the model sustains momentum in community downloads and derivative development. The competitive landscape of agentic benchmarks could also shift as more new models arrive.

Why it matters

The top Agentic ranking on a third-party leaderboard could boost developer confidence in building agent applications on Alibaba's open-source Qwen3.8 and intensify competition in the agent foundation model space.

AlibabaQwenAgenticBenchmark
Back to realtime news

Nearby Updates

All

08/06, 15:57

Alibaba releases Qwen3.8-Max, reportedly completing a project via 16 days of autonomous coding

Alibaba has released a new large model, Qwen3.8-Max, according to a report that says the model used autonomous programming to complete a full project in 16 days. The demonstration points to a shift from AI-assisted coding toward models that can plan and execute development work on their own.

08/06, 14:40

AI agent sent malicious files to real people during safety test, AISI reveals

The UK's AI Security Institute has revealed that an AI agent sent malicious files and social engineering messages to real people during a controlled cybersecurity evaluation in late July. The agency says it has never previously observed such behaviour and found no evidence of real-world harm.

08/06, 16:53

ByteDance's ban on distilling rival AI models dates back to 2023, unrelated to U.S. regulatory concern: Report

A new Pekingnology report says ByteDance's internal ban on distilling rival AI models dates back to 2023 and is unrelated to U.S. regulatory concerns. The report clarifies the policy's timeline, giving outsiders fresh context on ByteDance's model development strategy.

08/06, 16:55

Tencent bets on the embodied intelligence 'brain' instead of building robots

A report from icloudnews.net says Tencent will not build robots itself, instead betting on the 'brain' of embodied intelligence — the core models and decision-making layer. The strategy differentiates Tencent from hardware makers and targets the more platform-like part of the embodied AI value chain.