Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

Microsoft Internally Called AI Scraping 'the Largest Theft of Labor in Human History,' Unsealed Filings Show

Newly unsealed court filings show Microsoft privately described OpenAI's data practices as "theft," even as both companies scraped paywalled Times content and built datasets from it. The documents also record internal warnings that the approach would gut publishers.

Published
微软内部称 AI 抓取是“人类史上最大规模的劳动窃取”,解封文件曝光
Image source: techcrunch.com

A set of newly unsealed court filings has exposed, in unusually direct language, what Microsoft said about AI training data when it thought the public was not listening. According to TechCrunch, the unredacted documents show the company internally used the word "theft" to describe OpenAI's data practices.

The filings lay out three linked facts: both companies scraped paywalled Times content, used that material to build training datasets, and warned internally that the practice would gut publishers. The gap between the private assessment and the public posture is the sharpest thing in the record.

The most quoted line goes further than a legal complaint. AI scraping was described internally as "the largest theft of labor in human history," a framing that moves the argument from copyright compliance toward labor and value distribution, and shows the risk was understood inside the company rather than discovered later.

For the litigation, documents like these matter because they speak to knowledge and intent. Internal warnings are close to ideal evidence on that point, because a publisher does not have to prove scraping was technically possible, only that the defendant understood what it was doing and could foresee the harm.

The Times detail matters for a second reason. Paywalls are the core of many news organizations' business model, and if content behind them was folded into training data, the bargaining position of publishers in licensing talks narrows considerably.

The broader pattern is familiar: AI developers and publishers have spent years operating in a gray zone of license later, comply later. What the filings suggest is that the trade-offs were not invisible inside these companies. They were priced in and accepted.

Watch two things next. First, how widely other pending copyright cases cite these documents. Second, whether Microsoft and OpenAI adjust their public language in response. Either way, the training-data problem is getting harder to describe as a question of technical neutrality.

Why it matters

The filings turn a technical debate about training data into a question of intent, giving publishers stronger evidence in ongoing copyright suits and raising the pressure on AI companies at the licensing table.

MicrosoftOpenAICopyrightPolicy
Back to realtime news

Nearby Updates

All

09/18, 03:44

Anthropic Says Claude Optimized More Than 30 Open-Source Biomolecular Models

Anthropic reports that Claude has optimized more than 30 open-source biomolecular models, according to coverage from Unite.AI. The result extends general-purpose model capabilities deeper into life-science research workflows and revives the question of how much AI can accelerate discovery.

09/18, 04:00

UN teams with Google to make its global statistics readable by AI agents

The United Nations announced on September 17 that it is working with Google to launch the UN System Data Commons, a natural-language search layer built on Google's open-source Data Commons that supports the Model Context Protocol so AI systems can query UN statistics directly. The move follows a UNICEF benchmark in which six leading models averaged just 21.2% accuracy on global development questions.

09/18, 04:05

Report: US government website used a Chinese AI search tool the FBI said copied Anthropic

According to whbl.com, a US government website used an AI search tool from China that the FBI has said copied Anthropic's technology. The report links two threads usually discussed separately: allegations that Chinese AI products copied American models, and whether such tools are already running on public-facing government services.

09/18, 03:01

Spain's AI agent data breach report awaits AEPD review

eWeek reports that a data breach report involving an AI agent is now awaiting review by Spain's data protection authority, the AEPD. The case spotlights how regulators will draw compliance lines around agents that move data across systems.