Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

New benchmark PaperBenchX: models crack a millennium math problem, yet paper replication sits at 13.98%

According to Chinese outlet QbitAI, a new benchmark called PaperBenchX is being used to gauge AI's research capability: models can crack a millennium math problem, yet their rate of reproducing papers is as low as 13.98%. The result points to clear gaps in replicability on real research tasks.

Published
新基准PaperBenchX:模型能攻千禧年数学难题,论文复现率却低至13.98%
Image source: qbitai.com

According to Chinese outlet QbitAI, a new benchmark called PaperBenchX is being used to assess AI's research capability. The report says models can crack a millennium math problem, yet their success rate at reproducing papers is as low as 13.98%.

That 13.98% figure stands out because it separates solving problems from doing research. Math puzzles test reasoning and computation, while reproducing a paper requires a model to read a method, rebuild an experiment and arrive at verifiable results — much closer to real research.

The other focus of the report is the benchmark itself. Evaluations represented by PaperBenchX aim to provide a scale that better reflects research strength, shifting the question from whether a model can answer correctly to whether it can actually carry out research.

That distinction matters for the field. Competition scores and question-bank results have often been used to judge models, but such metrics can overstate performance on open-ended research tasks. A replication rate is closer to the problems researchers face day to day.

The report's wording leaves room for doubt: it suggests PaperBenchX may better measure AI's research strength, and the headline ends with a question mark, signaling an emerging evaluation idea rather than a settled verdict.

What to watch: the details and coverage of PaperBenchX, how model performance on replication changes across versions, and whether benchmarks like it become a new reference point for AI research capability. The report supplies a concrete number; the debate around it is just beginning.

Why it matters

PaperBenchX uses paper replication rate as a measure of AI research ability, and the low 13.98% score points to clear gaps on open-ended research tasks; benchmarks like it could change how the field evaluates models.

PaperBenchXBenchmarkAI Research
Back to realtime news

Nearby Updates

All