The Seattle Times and Newsday have filed a joint copyright infringement lawsuit against OpenAI and Microsoft, alleging that the technology companies systematically scraped their proprietary journalism to train generative AI models without authorization or compensation. The complaint, filed in the U.S. District Court for the Southern District of New York, marks a significant escalation in the ongoing legal conflict between traditional news publishers and the developers of large language models. The plaintiffs argue that their reporting, including content protected by subscription paywalls, has been ingested into systems like ChatGPT and Microsoft’s Copilot, effectively cannibalizing the traffic and advertising revenue that sustains their newsrooms.
This litigation joins a growing wave of legal challenges from media organizations seeking to establish whether the ingestion of copyrighted journalism for AI training constitutes fair use. The publishers are not merely seeking monetary damages; they are demanding the destruction of any copies of their work held by the defendants, as well as the removal of their content from the training datasets and the models themselves. This request for a total purge of their data from AI architectures represents an aggressive legal strategy that, if successful, could force a fundamental restructuring of how AI companies build and maintain their foundational models.
Allegations of Unauthorized Data Ingestion
The core of the lawsuit centers on the claim that OpenAI and Microsoft have built their commercial AI products on the back of decades of journalistic labor without obtaining licenses or providing payment. The publishers allege that these AI systems do not just learn from their work but frequently reproduce passages of their reporting verbatim in response to user queries. This behavior, the plaintiffs contend, often occurs without proper copyright attribution, which they argue compounds the harm to their intellectual property and brand integrity.
the complaint highlights a specific concern regarding the accuracy of AI-generated content. The newspapers allege that these models occasionally generate false information and incorrectly attribute it to the news organizations, creating potential reputational damage. By providing answers directly within the AI interface, the publishers argue that the defendants are effectively bypassing the need for users to visit the original news websites, thereby undermining the subscription-based business models that fund professional investigative journalism.
The Economic Stakes for Local Journalism
For regional publishers like The Seattle Times and Newsday, the stakes of this legal battle are existential. Alan Fisco, Chief Executive Officer of The Seattle Times, has framed the issue as a direct defense of the millions of dollars the newspaper invests annually in original reporting. The publishers argue that the current trajectory of AI development threatens to strip away the value of their work, leaving them unable to sustain the costs of gathering news if they cannot control how that information is monetized by third-party technology firms.
This lawsuit is particularly notable because it highlights the breakdown of the corporate-investment-as-partnership model. Microsoft, a defendant in the suit, has previously provided funding for journalism projects at The Seattle Times. The fact that a former partner is now a primary defendant underscores the deepening divide between the tech industry and the media sector. While some publishers, such as The Associated Press and Vox Media, have opted to negotiate licensing deals with OpenAI, the decision by these two regional outlets to pursue litigation suggests that many news organizations remain deeply skeptical of the current terms offered by AI developers.
A Growing Pattern of Legal Confrontation
The legal landscape for AI developers is becoming increasingly crowded as more publishers join the fray. The Seattle Times and Newsday are now part of a broader coalition of nearly 400 local newspapers that have recently initiated legal action against OpenAI and Microsoft. This case follows high-profile lawsuits from The New York Times, Ziff Davis, Merriam-Webster, and Encyclopedia Britannica, all of which are testing the limits of copyright law in the age of generative AI.
OpenAI has consistently maintained that its models are developed using publicly available information and that this practice is protected by fair-use principles. However, no court has yet issued a definitive ruling on whether the ingestion of copyrighted text for AI training falls under this legal doctrine. As these cases proceed through the court system, the tech industry faces the prospect of either being forced to pay significant retroactive damages or being required to fundamentally alter the datasets that power their most advanced models.
Unresolved Questions and Future Milestones
As the litigation moves forward, the industry is watching for any signs of a potential settlement or a landmark judicial decision. Microsoft has expressed surprise at the lawsuit while signaling an openness to discussing potential resolutions, suggesting that the company may be looking for a path to avoid a protracted legal battle. Conversely, the plaintiffs' demand for the destruction of models and datasets remains a major point of contention that could make a quick settlement difficult to achieve.
Until a court provides a clear ruling or Congress intervenes with new legislation, the media industry remains split between those pursuing litigation and those seeking licensing agreements. The outcome of this case will likely serve as a bellwether for the future of the information economy, determining whether AI companies can continue to operate under the current paradigm or if they will be required to treat human-created content as a licensed commodity.