Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

Bilibili launches AI arena with 100 models competing, GPT-6 on top

Bilibili launched its AI Infinite Arena on September 16 and published a first leaderboard in which GPT-6 Astra took the top spot across evaluations from ten creators, with domestic Chinese models taking three of the top five places, according to Securities Market Weekly. The board aggregates creator-run real-world tests of more than a hundred models and will update in real time.

Published
B站「AI无限竞技场」上线:全球百大模型同场竞技,GPT-6高居榜首
Image source: bilibili.com

Bilibili launched an AI Infinite Arena on September 16 and published the first edition of its model leaderboard at the same time. According to a report from Securities Market Weekly, GPT-6 Astra took first place across evaluations from ten Bilibili creators and recorded the most first-place finishes, ahead of GLM-5.3 on that count, while domestic Chinese models filled three of the top five spots.

The format is deliberately different from a research benchmark. The arena is a hub that collects model evaluations produced by Bilibili creators across more than a hundred models, built from real-world tests in categories such as coding, reasoning, collaboration and knowledge. Creators set their own tasks, from real workflows and professional applications to playful scenarios, and competing models face the same question.

The first leaderboard already spans DeepSeek, Kimi, ChatGPT, Claude, Gemini, Doubao, Qwen, Hy and MiniMax, which puts open and closed models, and Chinese and overseas systems, on the same table.

Bilibili's stated context is scale. Watch time for AI knowledge content on the platform grew 72% year over year, and more than 190 million users watch AI-related content each month. Live streams, videos and scrolling comments match how AI enthusiasts discover and debate information, and a large body of evaluation content has already accumulated inside the community. The arena organizes that material rather than building an evaluation system from scratch.

Notably, the ranking updates in real time and stays open to creators across the site. That makes it a recurring evaluation entry point rather than a one-off launch, with samples continuously replenished from real usage scenarios.

The value of any arena-style leaderboard depends heavily on how evaluations are sampled, who scores them and how results are aggregated. What has been disclosed so far stops at the layer of creators setting the task and models answering the same question; the full scoring and aggregation methodology has not been laid out. For model developers, board position is increasingly a public marketing asset, and the top spot most of all, which is exactly why transparency about the evaluation will draw scrutiny.

Three things to watch: whether Bilibili publishes a fuller methodology and complete ranking, whether the top models hold their positions once the board updates continuously, and whether a creator-driven evaluation becomes a common reference point for Chinese users judging how capable a model really is.

Why it matters

By letting community creators set the questions for a hundred models, Bilibili is doing with a content ecosystem what evaluation labs have long done with benchmarks, and a transparent, continuously updated board could become a mainstream reference point in China.

Bilibili大模型评测GPT-6AI Arena
Back to realtime news

Nearby Updates

All

09/16, 11:40

Hangzhou's Liwensuo opens Lévin Harness, an agent workspace for protein design

Hangzhou-based AI protein design company Liwensuo has released Lévin Harness, an agent-centred protein design application now open to the research community with Apple-silicon Mac support. It places data, models, plugins, compute and workflows in one workspace so that literature work, tool setup, GPU jobs and result analysis can run as a repeatable loop around the models.

09/16, 11:07

Hand the memory to the CPU: Intel lays out a data-center KV cache strategy

Intel has outlined a data-center approach that offloads the growing KV cache from GPU memory to CPU-side memory and storage so GPUs can focus on generating tokens, according to a QbitAI report published on September 16. Tests cited in the report show up to about 5x faster time to first token with tiered offloading and roughly 20% to 30% less cache space through lossless hardware compression.

09/16, 10:27

Huawei GTS teaches its ops agent to 'watch' networks troubleshoot, nearly clears dual-firewall test

Huawei's GTS unit has trained an agent to diagnose network faults by 'watching' the network, and it nearly cleared the hard dual-firewall scenario, according to qbitai. The reported results include a 24.2% lift in task pass rate and as much as 45% lower token cost.

09/16, 09:51

Ex-Anthropic researcher alleges the lab is accelerating the AI self-improvement race

A former Anthropic researcher has publicly alleged that the company is accelerating the race toward AI self-improvement, according to a report by chosun.com. The claim lands on a lab that has built its identity around safety, and so far the report offers the allegation itself rather than evidence outsiders can check.