Realtime AI News
Alibaba updates flagship Qwen3.8-Max, tops CodeArena front-end coding leaderboard
On September 2, Alibaba updated its flagship Qwen3.8-Max model with post-training focused on coding and professional office work, lifting it to first place on the CodeArena front-end WebDev leaderboard with a score of 1691. Alibaba says the new version averages about $5 per million tokens, and it is now live on the Qwen AI platform's API with Qwen Office, Qoder, and the Qwen app already integrated.
Alibaba updated its flagship model Qwen3.8-Max on September 2, promising significant performance gains over the previous version. The company describes Qwen3.8-Max as its most powerful large language model to date, with 2.4 trillion total parameters and 1 million tokens of context, and says the update adds specialized post-training for coding and professional office work.
The most visible result is on the leaderboards: on CodeArena, a third-party benchmark focused on front-end web development (WebDev), the new version gained 22 points to reach 1691, placing first overall ahead of models including Claude Opus 5 and Kimi K3.
Price is also part of the pitch. CodeArena's updated cost-performance ranking (the Pareto frontier) shows the new Qwen3.8-Max averaging about $5 per million tokens, undercutting every other model priced above that level.
Alibaba says the refreshed model shows stronger agentic coding abilities, making it better suited to complex real-world enterprise tasks, research, and long-horizon work.
The new model is now live on the Qwen AI platform with API services available, and Qwen Office, Qoder, and the Qwen app have all integrated it at launch.
Front-end coding has become one of the most crowded battlegrounds in the large-model race, and Alibaba is betting on a combination of a top leaderboard slot and aggressive pricing. The open question is whether a CodeArena score of 1691 translates into a real advantage in enterprise projects, and how rivals such as Claude Opus 5 and Kimi K3 respond.
Why it matters
Alibaba is using targeted post-training and aggressive per-token pricing to challenge pricier rivals at the top of coding benchmarks. Whether the CodeArena lead holds in real enterprise workloads, and how Claude Opus 5 and Kimi K3 respond, will determine if the position sticks.
Nearby Updates
All09/02, 14:10
Alibaba updates flagship Qwen3.8-Max, front-end coding tops global leaderboard
Alibaba has refreshed its flagship Qwen3.8-Max with post-training tuned for coding and professional office work, delivering a significant performance gain over the previous version. The new model scored 1691 on CodeArena, a leading front-end coding leaderboard, up 22 points, to place first overall ahead of models such as Claude Opus 5 and Kimi K3.
09/02, 14:05
OpenAI and Anthropic join the rush for Apple's Mac mini, AI's favorite hardware
OpenAI and Anthropic have joined the rush to buy Apple's Mac mini, making the desktop one of the AI industry's favorite pieces of hardware, according to Chinese financial outlet CLS. The report gives no purchase figures, but the move suggests even frontier labs are diversifying their compute beyond massive GPU clusters.
09/02, 13:57
Tencent's Hy4 preview returns it to the top tier of open-source AI, SCMP says
South China Morning Post says Tencent's Hy4 preview has put the company back in the top tier of open-source AI, calling the release a recommitment to the open-source path. The analysis notes that a credible open-weights model restores Tencent's competitiveness with developers and downstream adopters.
09/02, 13:33
OpenAI's new Astra model can build attacks without humans, reports say, as its highest safety mechanisms kick in
OpenAI has unveiled a new model, Astra, whose capability assessment reportedly exceeded the company's safety threshold and triggered its highest-level protection mechanisms. CoinDesk reports that OpenAI says Astra can build attacks autonomously, without human help.