Home Latest Insights | News Seattle Times, Newsday Sue OpenAI and Microsoft Over Use of News Articles to Train AI

Seattle Times, Newsday Sue OpenAI and Microsoft Over Use of News Articles to Train AI

Seattle Times, Newsday Sue OpenAI and Microsoft Over Use of News Articles to Train AI

The Seattle Times and Newsday sued OpenAI and Microsoft in federal court on Friday, accusing the technology companies of using their journalism without permission to train and operate artificial intelligence systems that can reproduce or closely mimic their reporting.

The lawsuit, filed in the U.S. District Court for the Southern District of New York, alleges that OpenAI and Microsoft scraped the newspapers’ websites, including material available only to paying subscribers, and incorporated their articles into datasets used to develop and operate products including ChatGPT, Microsoft Copilot and AI features within Bing.

The newspapers said the alleged use of their work goes beyond simply training AI models. They argued that the resulting products can reproduce passages from their articles, closely paraphrase their reporting and generate answers that give users information without requiring them to visit the publishers’ websites or purchase subscriptions.

That creates a potentially fundamental threat to the business model underpinning digital journalism, the newspapers said. Publishers spend heavily on reporters, editors, investigations and other newsgathering operations, while AI systems can potentially extract and redistribute the resulting information at scale.

“We feel strongly that we must defend our content – which we spend millions of dollars a year to produce – from being used without our consent or compensation,” Seattle Times President and CEO Alan Fisco wrote to employees, according to the newspaper.

An OpenAI spokesperson said the company’s models are trained on publicly available data and that their use of such material is protected by fair use. The spokesperson did not specifically comment on the lawsuit.

Microsoft, which is based near Seattle, said it was surprised by the legal action but acknowledged the importance of local journalism.

“While we’re surprised by the lawsuit, we appreciate the importance of local journalism and we’re always happy to sit down and explore solutions to this type of dispute,” a Microsoft spokesperson said in an email.

The newspapers are seeking an order requiring the companies to destroy copies of their copyrighted works as well as any training datasets or AI models that incorporate those works.

Such a remedy could have consequences well beyond the two publishers if the court ultimately finds that copyrighted news content was unlawfully incorporated into AI systems. Removing specific material from already trained models and datasets can be technically difficult, potentially turning a copyright dispute into a question about how AI companies should remediate models after they have been trained.

The case adds to a rapidly expanding legal confrontation between publishers and AI developers over who should control and benefit from the enormous amount of information used to build generative AI.

The New York Times filed a similar lawsuit against OpenAI and Microsoft in 2023, accusing the companies of using millions of its articles without authorization to develop AI systems. That case remains pending and has become one of the most closely watched copyright disputes in the technology industry.

Dozens of other copyright holders have also sued AI companies including OpenAI, Anthropic and Meta, alleging that their books, images, software, news articles and other creative works were used without permission to train AI models.

At the center of many of the cases is the question of whether training an AI model on copyrighted material constitutes a lawful use of that material, and whether AI-generated outputs that reproduce or closely substitute for original works create a separate copyright or economic harm.

The publishers’ argument is focused on that second issue. Even if courts ultimately permit some forms of data use for model training, publishers could argue that AI systems should not be allowed to reproduce substantial portions of their reporting or answer questions in ways that substitute for the original article.

That distinction could prove important for the future economics of online news. Search engines historically directed readers to publishers, creating a flow of traffic that could be monetized through advertising and subscriptions. AI assistants can instead provide synthesized answers directly, potentially reducing the incentive for users to click through to the source.

For local newspapers such as the Seattle Times and Newsday, the stakes have become high. Unlike large technology companies, publishers generally depend on subscription revenue, advertising, and audience engagement to finance expensive reporting operations. If AI systems capture the informational value of that reporting while reducing visits to the original publisher, the economic impact could extend beyond copyright royalties.

The lawsuit therefore puts pressure on OpenAI and Microsoft to address not only whether their use of news content is legally permissible, but also how AI companies should compensate publishers whose reporting contributes to the systems’ usefulness. The companies have already faced growing pressure to establish licensing arrangements and other commercial relationships with content owners. But litigation remains the more consequential route for publishers seeking to establish legal boundaries that could apply across the industry.

The outcome, adding to similar lawsuits, is expected to help determine whether the AI industry can continue relying broadly on internet content under existing copyright doctrines or whether developers will need permission, licensing agreements, or other compensation mechanisms to use professional journalism in building commercial AI systems.

For publishers, the issue has become about whether the economics of producing original information can survive when AI systems are capable of absorbing that information, repackaging it, and delivering it directly to consumers without sending those consumers back to the organizations that paid to produce it.

No posts to display

Post Comment

Please enter your comment!
Please enter your name here