Seattle Times and Newsday sue OpenAI and Microsoft for infringement
Original reporting by The Verge

A series of high-profile copyright infringement lawsuits are challenging the foundational practices of leading AI developers, notably OpenAI and Microsoft, over the alleged unauthorized use of published works for training generative AI models. This legal offensive represents a critical juncture for the burgeoning AI industry, as content creators grapple with how their intellectual property is being consumed and repurposed by advanced algorithms. The Seattle Times and Newsday are the latest prominent media outlets to file suit, accusing OpenAI of leveraging their journalistic content as training data without permission and subsequently reproducing passages in response to user queries. Their action echoes similar legal battles initiated by major entities like The New York Times, Ziff Davis, Merriam-Webster, and Encyclopedia Britannica, all alleging the misappropriation of copyrighted material to build powerful AI systems.
Escalating Demands These new filings also target Microsoft, a partner in OpenAI's Copilot technology, and notably join a class action lawsuit from nearly 400 local newspapers that have collectively sued both companies. Beyond direct infringement, publishers contend that AI chatbots, by providing distilled information sourced directly from their sites, diminish the need for direct visits, thereby eroding crucial subscription and advertising revenue essential for sustaining journalism. Crucially, The Seattle Times and Newsday are seeking not just monetary damages but the unprecedented destruction of their copyrighted works, all associated training datasets, and any AI models that have incorporated their content, marking a significant escalation in the ongoing legal battle over AI's impact on intellectual property and the future of content creation.
The escalating legal actions, spearheaded by venerable institutions like The Seattle Times and Newsday, against AI giants OpenAI and Microsoft represent a watershed moment in the evolving relationship between generative AI and original content creation. These lawsuits move beyond mere compensation claims, directly challenging the foundational methodologies of large language models by demanding the destruction of infringing datasets and models. The core argument—that unauthorized use of copyrighted material for training constitutes infringement and undermines the economic viability of journalism and other creative industries—strikes at the heart of AI’s current development paradigm, forcing a reckoning with the definition of ‘fair use’ in the digital age.
Reshaping AI's Foundation
The outcomes of these cases will have profound implications, not just for the defendants, but for the entire AI industry. A ruling in favor of the publishers could necessitate a radical rethinking of how AI models are trained, potentially ushering in an era where data acquisition is strictly licensed and compensated, rather than broadly scraped from the open web. Such a shift would inevitably impact the diversity and scope of future AI models, possibly favoring those with deep pockets for data licensing or those relying exclusively on public domain and ethically sourced proprietary datasets. For content creators, particularly news organizations facing existential threats, these proceedings could establish critical precedents, securing a more equitable framework for their intellectual property in the age of artificial intelligence and potentially redirecting significant revenue streams back to the originators of information, ensuring the continued production of high-quality, verifiable content.
Frequently asked questions
- What are news publishers suing AI companies like OpenAI and Microsoft for?
- News publishers allege that AI companies like OpenAI and Microsoft unlawfully used their copyrighted journalism as training data for large language models. They claim this constitutes copyright infringement because the AI models then reproduce passages or provide answers that reduce the need for users to visit the original news sites, impacting subscription revenue. Plaintiffs are seeking compensation and the destruction of models incorporating their work.
- Why do media outlets want AI models built on their content to be destroyed?
- Media outlets are demanding the destruction of AI models because they believe these models incorporate their copyrighted works without permission. They argue that if the models were trained on infringed content, the models themselves are tainted. Destroying these models and associated datasets is seen as a way to prevent continued unauthorized use and safeguard their intellectual property rights and business models.
- Which AI companies are currently facing copyright infringement lawsuits from news organizations?
- OpenAI and Microsoft are prominent AI companies currently facing numerous copyright infringement lawsuits from news organizations. Publishers such as The Seattle Times, Newsday, and The New York Times allege that these companies used their copyrighted articles to train AI models like Copilot without authorization. These lawsuits highlight ongoing legal challenges regarding intellectual property rights in the age of generative AI.