News Publishers Accuse Microsoft and OpenAI of Using News Content Without Proper Licensing
News organizations are accusing Microsoft and OpenAI of ignoring warnings that scraping news content could amount to theft. In a court filing, the publishers alleged that the companies chose not to license the content and instead obtained or reused it through other channels.
Microsoft Allegedly Resold Bing Data for AI Training
The motion accuses Microsoft of violating “industry norms” by selling datasets it purchased for Bing as training data for OpenAI. According to the news organizations, Microsoft allegedly did so without consulting publishers that would not have approved the reuse of their content under basic search-engine crawling consent.
OpenAI Accused of Using a 1.8 Million-Article Dataset
The publishers also alleged that OpenAI acquired a New York Times dataset containing 1.8 million articles from a third party that was bound by an agreement prohibiting commercial use of the data.
According to the news organizations, OpenAI employees knew it was “not appropriate” to use the dataset to “train models,” but used it anyway. The allegations were included in the publishers’ motion and have not been established as fact.
Publishers Say AI Companies Are “Free Riding” on News
The motion argued that the dispute extends beyond Microsoft and OpenAI to other AI companies that allegedly “free ride” on news publishers’ content. The publishers specifically pointed to Google’s AI Overview, which was introduced soon after ChatGPT and began absorbing traffic that previously went to news websites.
The news plaintiffs argued that both publishers and AI companies could face serious consequences unless courts clarify whether AI companies must license news content used to develop their systems.
Microsoft Documents Address AI’s Impact on Content Creators
The publishers cited an internal Microsoft document that acknowledged a “real risk” that generative AI could “significantly disrupt” the employment of “the very people who generated the data used to train the underlying models.” They also referenced a Microsoft comic that illustrated how large language models could disrupt the industry’s supply chain.
Credit:
via News Plaintiff
The Publishers’ “Loop of Destruction” Argument
“AI companies are still powerless to break out of this ‘loop of destruction,’ because while the industry as a whole would benefit if each company paid to maintain the continued production of creative works that rely on its technology, individual companies would be better off getting content for free while others paid,” the news organizations argued.
As evidence of what they described as this “blind greed,” the motion highlighted that Brockman was “deeply motivated by the hundreds of millions in profits” that could be made by commercializing OpenAI’s technology.
“If it turns out that copying news for AI is not fair use, this Prisoner’s Dilemma will be resolved by leveling the playing field for all AI companies, including OpenAI and Microsoft,” the news organizations said.
This article has been updated with quotes from New York Daily News consultant Stephen Lieberman.
Source: arstechnica.com


