Realtime AI News
New benchmark PaperBenchX: models crack a millennium math problem, yet paper replication sits at 13.98%
According to Chinese outlet QbitAI, a new benchmark called PaperBenchX is being used to gauge AI's research capability: models can crack a millennium math problem, yet their rate of reproducing papers is as low as 13.98%. The result points to clear gaps in replicability on real research tasks.
According to Chinese outlet QbitAI, a new benchmark called PaperBenchX is being used to assess AI's research capability. The report says models can crack a millennium math problem, yet their success rate at reproducing papers is as low as 13.98%.
That 13.98% figure stands out because it separates solving problems from doing research. Math puzzles test reasoning and computation, while reproducing a paper requires a model to read a method, rebuild an experiment and arrive at verifiable results — much closer to real research.
The other focus of the report is the benchmark itself. Evaluations represented by PaperBenchX aim to provide a scale that better reflects research strength, shifting the question from whether a model can answer correctly to whether it can actually carry out research.
That distinction matters for the field. Competition scores and question-bank results have often been used to judge models, but such metrics can overstate performance on open-ended research tasks. A replication rate is closer to the problems researchers face day to day.
The report's wording leaves room for doubt: it suggests PaperBenchX may better measure AI's research strength, and the headline ends with a question mark, signaling an emerging evaluation idea rather than a settled verdict.
What to watch: the details and coverage of PaperBenchX, how model performance on replication changes across versions, and whether benchmarks like it become a new reference point for AI research capability. The report supplies a concrete number; the debate around it is just beginning.
Why it matters
PaperBenchX uses paper replication rate as a measure of AI research ability, and the low 13.98% score points to clear gaps on open-ended research tasks; benchmarks like it could change how the field evaluates models.
Nearby Updates
All10/08, 18:23
Lenovo opens blind pre-orders for YOGA Pro 15 with Nvidia RTX Spark N1X chip
Lenovo has opened blind pre-orders for its YOGA Pro 15 RTX Spark laptop, positioned around Nvidia's RTX Spark N1X chip. Full pricing and detailed specs are expected to be revealed later.
10/08, 18:37
India's minister says the time for AI rules has come, with a consultation paper due within a month
A report says India's minister Ashwini Vaishnaw has declared that the time has come for AI regulations, with the central government set to issue a consultation paper within a month. The statement signals that India is moving toward a concrete AI governance framework.
10/08, 19:40
JD.com teams with partners to push 'agent computers' with a 100 billion yuan two-year sales goal
JD.com is joining industry chain partners to drive growth in a new 'agent computer' category, targeting more than 100 billion yuan in sales over two years. The move extends AI from software into a new hardware category defined jointly by channel and vendors.
10/08, 14:57
OpenAI unveils 722 AI math solutions, sparking community discussion
OpenAI has unveiled 722 AI math solutions, according to reports aggregated by Google News, a release that became a talking point. The report says the batch has prompted community discussion about human understanding of the underlying problems.