Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

UN teams with Google to make its global statistics readable by AI agents

The United Nations announced on September 17 that it is working with Google to launch the UN System Data Commons, a natural-language search layer built on Google's open-source Data Commons that supports the Model Context Protocol so AI systems can query UN statistics directly. The move follows a UNICEF benchmark in which six leading models averaged just 21.2% accuracy on global development questions.

Published
联合国联手谷歌上线 Data Commons:把全球统计数据做成 AI 可直接调用的底座
Image source: techcrunch.com

The United Nations said on Thursday that it is working with Google to rebuild its collection of global statistics so AI systems can find and use it reliably. The new platform, called the UN System Data Commons, is built on Google's open-source Data Commons and lets users search statistics from across UN agencies with natural-language queries, replacing the older UNData portal where people largely had to browse a traditional database interface.

The detail that matters most to developers is support for the Model Context Protocol, the standard that lets AI systems connect directly to external data sources. Instead of relying on training data that may be stale, an agent can pull official UN statistics on demand, and the platform tracks where each number came from so an AI-generated answer can be traced back to its original source.

The UN's push follows evidence that models handle this kind of data poorly. João Pedro Azevedo, UNICEF's chief statistician, told reporters that a UNICEF benchmark of six large language models across more than 133,000 responses on global development indicators produced an average accuracy score of just 21.2%. The test covered OpenAI's GPT-4o and GPT-4o-mini, Anthropic's Claude Sonnet 4.5 and Haiku 4.5, and Google's Gemini 2.5 Flash and Gemini 2.0 Flash.

Consistency was as much a problem as accuracy. About three in five responses did not provide a usable number at all, often because the models hedged their answers, Azevedo said. When the same questions were run again on the same model versions roughly two days later, models that produced a number both times returned the identical value only about half the time. The study is a UNICEF working paper being prepared for journal submission and has not been peer-reviewed; the organisation plans to release its methodology, code and data alongside the paper.

Demand from AI assistants is already visible in the traffic. UNICEF's data website receives more than six million visits a month, Azevedo said, and referrals from links inside ChatGPT answers rose 67% year over year between January 1 and September 14, accounting for 6.4% of all sessions this year; AI assistants overall now drive roughly one in ten visits.

On coverage, the UN said 26 of its entities have committed to the Data Commons, with data from nearly 20 available at launch, and it aims to bring 80% of the UN system's statistical datasets onto the platform by 2027. Google.org provided $2 million in capacity-building funding and technical support for the core infrastructure. Prem Ramaswamy, who leads Google's Data Commons team, said the system is hosted on a UN-governed instance and is intended to be maintained, operated and scaled independently by the UN.

A demonstration showed what that unlocks: connected through MCP, an AI system asked about the impact of the U.S. President's Emergency Plan for AIDS Relief in Africa pulled together UN statistics on HIV infections, AIDS mortality and life expectancy and produced an infographic, without a user manually assembling the underlying datasets.

Ramaswamy added a caveat that applies to every data-to-agent project: authoritative inputs do not guarantee authoritative conclusions. "Because models can misinterpret nuance, a human should always review the outputs before citing or publishing them," he said. For the UN, the platform is both a fix for AI accuracy and a play for visibility — if assistants become the default way people reach development data, being machine-readable with traceable sourcing is how the institution stays in the answer.

Why it matters

If MCP becomes the default interface between public data and AI, official statistics stop being something models guess at and become something agents call, shifting the accuracy problem from the model side to the data side.

GoogleUNMCPData
Back to realtime news

Nearby Updates

All

09/18, 04:05

Report: US government website used a Chinese AI search tool the FBI said copied Anthropic

According to whbl.com, a US government website used an AI search tool from China that the FBI has said copied Anthropic's technology. The report links two threads usually discussed separately: allegations that Chinese AI products copied American models, and whether such tools are already running on public-facing government services.

09/18, 03:46

Microsoft Internally Called AI Scraping 'the Largest Theft of Labor in Human History,' Unsealed Filings Show

Newly unsealed court filings show Microsoft privately described OpenAI's data practices as "theft," even as both companies scraped paywalled Times content and built datasets from it. The documents also record internal warnings that the approach would gut publishers.

09/18, 03:44

Anthropic Says Claude Optimized More Than 30 Open-Source Biomolecular Models

Anthropic reports that Claude has optimized more than 30 open-source biomolecular models, according to coverage from Unite.AI. The result extends general-purpose model capabilities deeper into life-science research workflows and revives the question of how much AI can accelerate discovery.

09/18, 03:01

Spain's AI agent data breach report awaits AEPD review

eWeek reports that a data breach report involving an AI agent is now awaiting review by Spain's data protection authority, the AEPD. The case spotlights how regulators will draw compliance lines around agents that move data across systems.