‘Largest theft in human history’: OpenAI accused of using millions of news articles

New York: OpenAI committed an “astonishing theft of unprecedented proportions” by using millions of news articles to train its artificial intelligence models, The New York Times has alleged in a court filing unsealed on Thursday.
The San Francisco-based AI company is accused of scraping more than 10 million articles, with nearly one-third of them reportedly coming from The New York Times.
The filing cites comments allegedly made by Microsoft Director of Applied Science Brent Hect, who described the use of the articles as an “astonishing theft of unprecedented proportions”. He also reportedly called it potentially the “largest theft of labour in human history”.
Hect is further alleged to have warned that OpenAI may have engaged in an “accidental cover up” while attempting to identify material from The New York Times and other plaintiffs within its systems.
Microsoft, however, said Hect's comments represented his personal views and did not reflect the company's position.
A Microsoft spokesperson told AFP that the statements were the “individual perspective” of one employee and “do not represent the company's views”.
The New York Times sued OpenAI and Microsoft in a New York federal court three years ago, accusing the companies of using its copyrighted material without permission to train OpenAI's ChatGPT models. Microsoft has been an investor in OpenAI since 2019.
Several other publishers have joined the legal action, including Ziff Davis, which owns CNET and other technology publications, as well as the parent company of Mother Jones, The Intercept and several local newspapers in the US.
The publishers are seeking damages for each article they allege was copied and used by OpenAI's models. The potential total damages have not been determined.
OpenAI has since entered into licensing agreements with several news organisations around the world. However, concerns remain among publishers that generative AI services could reduce traffic to news websites by providing information directly to users.
AI companies have attempted to address the issue by displaying citations and links alongside answers. The court filing, however, claims that even OpenAI employees questioned whether such links would drive traffic.
“One of OpenAI's own engineers admitted, according to the filing, that users would not click on the links regardless of how prominently they were displayed.”
OpenAI and Microsoft maintain that their use of news content is transformative and protected under US fair-use laws.
Earlier this month, the US Department of Justice filed a brief supporting OpenAI and Microsoft, citing scientific progress, economic growth and national security among its arguments.
The publishers have asked the court to grant summary judgment in their favour.
If US District Judge Sidney Stein grants the request, the case would be decided without going to trial. A ruling is not expected until 2027.