The Seattle Times and Newsday have initiated a federal copyright lawsuit against OpenAI and Microsoft, alleging that the technology giants systematically scraped their journalism without permission to train commercial artificial intelligence products. Filed on Friday in the U.S. District Court for the Southern District of New York, the legal action targets the incorporation of paywalled and open web content into training datasets for systems such as ChatGPT, Microsoft Copilot, and Bing's AI features.
The fresh litigation marks a significant escalation in the ongoing legal confrontations between media publishers and major AI developers over data ingestion practices. The newspapers argue that the resulting AI tools are capable of reproducing verbatim passages or closely paraphrasing proprietary reporting, thereby deflecting web traffic and undermining reader subscription revenues.
Allegations of Unauthorized Scraping and Paywall Evasion
The complaint asserts that OpenAI and Microsoft bypassed digital barriers, including paywalls, to harvest extensive archives of news content. According to the court filing, this harvested journalism served as foundational training data for prominent generative models and chatbot interfaces. The publishers contend that such automated collection infringes upon established copyright protections, taking commercial value without compensating the organizations that fund original reporting.
Executives at the affected publications emphasized the high financial stakes involved in maintaining investigative desks and regional newsrooms. Seattle Times President and CEO Alan Fisco addressed the internal workforce regarding the decision to litigate. "We feel strongly that we must defend our content - which we spend millions of dollars a year to produce - from being used without our consent or compensation," Fisco wrote, according to the newspaper.
Impact on Subscriptions and Reader Traffic
Beyond the initial ingestion of articles, the lawsuit highlights the downstream market harm caused by generative search and chat features. The plaintiffs argue that when systems like Microsoft Copilot or ChatGPT provide comprehensive answers derived from their reporting, end users have little incentive to visit the originating websites. This reduction in direct traffic directly threatens digital advertising yields and subscription conversions.
Local journalism operations depend heavily on reader acquisition funnels that convert search engine referrals and direct visits into paid subscribers. By synthesizing reporting into direct answers, the tech platforms allegedly appropriate the economic value of the newsroom's labor. The lawsuit asks the court to mandate a strict remedy: the complete destruction of copies of the newspapers' works, alongside any training datasets or AI models that incorporate their material.
Corporate Responses and Defense Frameworks
Defendants OpenAI and Microsoft offered differing initial reactions following the public filing of the complaint. San Francisco-based OpenAI defended its foundational practices through a spokesperson, stating that its models are trained using publicly available data and remain firmly grounded in fair-use principles, though the company declined to address the specific allegations in the new lawsuit.
Microsoft expressed surprise at the escalation while acknowledging the vital democratic function served by local news organizations. A Microsoft spokesperson stated in an email: "While we’re surprised by the lawsuit, we appreciate the importance of local journalism and we’re always happy to sit down and explore solutions to this type of dispute." Despite openness to dialogue, the software giant faces mounting legal scrutiny regarding how its ecosystem integrates harvested data.
Precedent in the New York Times Litigation and Industry-Wide Scrutiny
The action filed by The Seattle Times and Newsday does not exist in a vacuum, closely echoing a landmark 2003-initiated legal battle launched by The New York Times against the same corporate defendants. That ongoing federal case similarly accuses OpenAI and Microsoft of utilizing millions of unauthorized newspaper articles to train conversational models.
Copyright holders across the media, publishing, and creative sectors have increasingly turned to the federal judiciary to contest generative AI training methodologies. Similar copyright infringement lawsuits have been brought against other major artificial intelligence developers, including Anthropic and Meta. These coordinated legal challenges seek to establish firm judicial boundaries governing fair use, data scraping, and corporate compensation for intellectual property in the digital age.