Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

Alibaba updates flagship Qwen3.8-Max, front-end coding tops global leaderboard

Alibaba has refreshed its flagship Qwen3.8-Max with post-training tuned for coding and professional office work, delivering a significant performance gain over the previous version. The new model scored 1691 on CodeArena, a leading front-end coding leaderboard, up 22 points, to place first overall ahead of models such as Claude Opus 5 and Kimi K3.

Published

On September 2, Alibaba updated its flagship large language model Qwen3.8-Max, saying the new version delivers significantly stronger performance than its predecessor. The model is already available via API on the Qianwen AI platform, with integrations live in Qianwen Office, Qoder, and the Qianwen app.

The refresh centers on specialized post-training for coding and professional office work (Cowork), which Alibaba says boosts overall capability. The company highlights stronger agentic coding skills suited to complex real-world enterprise tasks, scientific research, and long-horizon work.

On CodeArena, a widely cited third-party ranking focused on front-end (WebDev) coding ability, the new Qwen3.8-Max gained 22 points to reach 1691, placing first overall ahead of models such as Claude Opus 5 and Kimi K3.

Cost is another differentiator. CodeArena's updated cost-performance (Pareto frontier) list shows the new Qwen3.8-Max averaging only about $5 per million tokens, undercutting every competing model priced above $5 per million tokens.

Qwen3.8-Max is the most powerful large language model Alibaba's Qwen line has built to date, with 2.4 trillion total parameters and support for a 1-million-token context window. That combination underpins its handling of complex coding and office workloads.

The rivals it leapfrogged, Claude Opus 5 and Kimi K3, are among the strongest coding models on the chart, so taking first place signals that Alibaba's flagship now competes head-to-head with leading international models in front-end and agentic coding.

The next thing to watch is real-world feedback from enterprise deployments, and whether Alibaba keeps up this cadence and extends the generation's strengths across more coding and office benchmarks.

Why it matters

The CodeArena No.1 spot puts Alibaba's flagship in the global front-end and agentic coding race, and a roughly $5-per-million-token price point adds serious cost pressure for pricier rivals.

AlibabaQwen3.8-MaxModel Update
Back to AI Daily

Nearby Updates

All

09/02, 14:10

Alibaba updates flagship Qwen3.8-Max, tops CodeArena front-end coding leaderboard

On September 2, Alibaba updated its flagship Qwen3.8-Max model with post-training focused on coding and professional office work, lifting it to first place on the CodeArena front-end WebDev leaderboard with a score of 1691. Alibaba says the new version averages about $5 per million tokens, and it is now live on the Qwen AI platform's API with Qwen Office, Qoder, and the Qwen app already integrated.

09/02, 14:06

Former ByteDance RL expert Sun Peng joins Physical AI firm 星尘智能 to round out its stack

QbitAI reports that Sun Peng, a former reinforcement-learning expert at ByteDance and former head of the agent center at Tencent Robotics X, formally joined 星尘智能 on September 2. The hire is aimed at completing the company's full-stack technology layout in Physical AI.

09/02, 14:05

OpenAI and Anthropic join the rush for Apple's Mac mini, AI's favorite hardware

OpenAI and Anthropic have joined the rush to buy Apple's Mac mini, making the desktop one of the AI industry's favorite pieces of hardware, according to Chinese financial outlet CLS. The report gives no purchase figures, but the move suggests even frontier labs are diversifying their compute beyond massive GPU clusters.

09/02, 14:20

Ant Group's OmniTable wins VLDB 2026 industrial best paper, handling 35PB of LLM corpus

QbitAI reports that Ant Group's unified wide-table system, OmniTable, has won the best-paper award in the industrial track at VLDB 2026. Built for large-model data preparation, the system handles a 35PB corpus on a single wide-table architecture and reportedly lifts processing efficiency by 5.6x.