Realtime AI News
Bilibili launches AI Infinite Arena leaderboard with 100 models competing
Chinese tech coverage reports that Bilibili has launched an evaluation leaderboard called AI Infinite Arena, describing it as one competition in which one hundred models from around the world take part. It marks the video platform's first move into ranking models, giving Chinese developers and creators a community-owned reference point beyond vendor self-reports and overseas arenas.

Chinese tech coverage reports that Bilibili has launched an AI evaluation leaderboard called AI Infinite Arena, describing it as a single arena in which one hundred models from around the world compete. The launch is a platform-level product move: it pulls model comparison out of research groups, overseas arenas and vendor-written model cards and into Bilibili's own community ecosystem.
The information available so far is narrow. The confirmed points are that the leaderboard is live and that it is framed around roughly a hundred models competing together. How the evaluation is designed, which models and which versions are admitted, and how often the ranking refreshes are not spelled out in the headline and summary. That makes this story about who runs the evaluation rather than about any single new score.
That distinction matters because leaderboards have become the default first filter for model selection. Developers comparing options routinely start from a ranking and then narrow down, so the measurement method quietly decides which capabilities look strongest. When rankings live on a platform's own site, the question of measurement becomes both a research question and a product question.
Bilibili's position makes the move notable. It is one of China's largest video communities for younger users and a dense home for technical, coding and AI-related content. A model arena hands that audience a shared comparison point, and hands the platform a topic that generates debate, side-by-side tests and derivative videos, all of which feed its content loop.
Model makers have a straightforward incentive as well: an extra leaderboard is an extra distribution channel for their names. For platforms, keeping evaluation on-site turns user curiosity about model capability into longer community sessions. Those two pressures help explain why platform companies increasingly run their own benchmarks instead of only citing third-party results.
What to watch next is the documentation. Whether Bilibili publishes its judging method, how it verifies model versions, whether it covers Chinese-language and multimodal tasks, and whether the board keeps updating will decide if this becomes a durable reference or a one-off talking point. A ranking without a transparent method can still travel widely as news, but it cannot yet be used as evidence.
Why it matters
China's model ecosystem gains another platform-owned evaluation entry point and model teams gain another channel for visibility, but the ranking only becomes a usable reference once its method is published.
Nearby Updates
All09/20, 14:48
Former OpenAI Researcher Releases Jev, a Model for Fast, Structured Software Decisions
According to OSCHINA, a former OpenAI researcher has released a model called Jev that aims to help software make fast, structured decisions. It is another attempt to pull decision-making out of general-purpose chat models, though public details currently stop at the announcement itself.
09/20, 13:05
Researchers Accessed OpenAI Systems via Claude AI in 72 Hours, Report Says
Researchers reportedly accessed OpenAI systems using Anthropic’s Claude AI within 72 hours, according to a report from Chosun Ilbo. The 72-hour window is the most striking detail, hinting at how quickly frontier models can be applied to probing large online systems, though the report says little about scope, authorization or impact.
09/20, 13:03
Kuakua Jingling Launches Two Revenue-Boosting AI Agents, Betting on the “AI + Digital Employee” Track
Kuakua Jingling has released two AI agents positioned around revenue generation, framing the launch as a step into what it calls the “AI + digital employee” track for enterprise AI, according to a report from Sina Mobile. The move signals a shift in enterprise AI competition away from generic model capability and toward agents expected to deliver measurable business results.
09/20, 12:50
Anthropic taps Accenture to audit frontier models as the consultancy takes on an AI oversight role
Accenture is taking on a new AI oversight role, with Anthropic tapping the consultancy to audit its frontier models, according to AD HOC NEWS. The key point is that a frontier lab is bringing an outside firm into model evaluation instead of keeping safety review entirely in house.