Seattle Times and Newsday Sue OpenAI Over Training Data
Two more newspapers join the copyright fight against OpenAI and Microsoft, sharpening questions about fair use, licensing, and who profits from journalism.
Written by AI. Marcus Chen-Ramirez

The Seattle Times and Newsday have joined the expanding group of newspapers suing OpenAI and Microsoft over the use of journalism in AI training, according to techbuzz.ai. Two plaintiffs is not a wave by itself. But these two are joining a docket that already includes some of the biggest names in American news, and each new filing makes the defendants' central argument, that the suits are isolated complaints rather than a structural dispute, harder to sustain.
What the New Lawsuits Claim
The techbuzz.ai report frames the dispute around two linked questions: whether model developers may ingest copyrighted news archives without permission, and whether commercial reuse of that material creates liability even when a model never reproduces an article verbatim. That second half is where the litigation gets interesting, because it moves the argument away from the familiar piracy framing (did the model output my exact paragraphs?) toward something closer to a market-substitution claim: the model absorbed decades of local reporting and now answers questions readers used to answer by clicking.
Both plaintiffs fit a pattern. The Seattle Times is a regional metro daily with deep local archives; Newsday covers Long Island with the same kind of decades-spanning record. Neither owns the global brand recognition of the New York Times, which filed the landmark suit against OpenAI and Microsoft in December 2023. That is arguably the point. A model that can answer "what happened with that zoning fight in Bellevue in 2019" has ingested the granular, expensive, locally sourced reporting that regional papers produce and that national AI companies did not pay for.
The Defendant's Case, Stated at Its Strongest
OpenAI has been public about its position. In its official response to the New York Times litigation, posted at OpenAI's own account of the case, the company argues that the lawsuit misrepresents how its models work, that training is legitimate, and that its use of copyrighted works constitutes "fair use." The company has also pointed out that the Times's engineering experts, in OpenAI's telling, had to manipulate prompts to elicit anything resembling regurgitation, and that the Times had previously collaborated with OpenAI through licensing deals, including one with the Associated Press that the Times invested in.
There are real arguments behind that posture. American fair-use doctrine has long tolerated technologies that copy works in order to do something new with them: search engines scanning the web, VCRs recording broadcasts, snippet-level quotation in search results. Without some form of the fair-use exception applied to machine learning, American AI developers face a permission-requirements regime that their European and Chinese competitors may navigate differently, or not at all. OpenAI's defenders make a national-competitiveness argument as much as a legal one.
The strongest version runs something like this: training a model is transformative because the model learns statistical patterns, not passages; the output does not compete with the articles consumed; and society gets a useful technology as a byproduct of otherwise lawful learning. Whether a court buys any part of this is unsettled.
The Publishers' Case, Stated at Its Strongest
The publishers' answer is that transformation doctrine was built for technologies that added value on top of works, not ones that could replace them. A search engine sends traffic to the original article; an answer engine delivers the answer and keeps the click. When a model can summarize or respond from an archive of reporting, the publisher loses the reader, the ad impression, and the subscription conversion, all while the model provider commercializes the extracted value.
This is the practical issue the techbuzz.ai piece flags directly: publishers are challenging not only training practices but the value extracted when models summarize, answer from, or substitute for original reporting. That distinction matters for remedy, too. Even if a court accepted that training itself is fair use, ongoing output that competes with the source material could still infringe, because fair use is a case-by-case weighing, not a blanket immunity.
The New York Times litigation crystallized this fight. A Harvard Law Review analysis of the case examined the competing arguments over whether model training is "transformative," noting the tension between theTimes's claim that OpenAI's products substitute for its journalism and OpenAI's claim that learning from copyrighted text at scale is fundamentally different from republishing it. That tension is exactly what the Seattle Times and Newsday filings extend: two more evidence sets showing model behavior against regional archives, two more courts weighing substitution.
What the Government Has Said
Courts will not be writing doctrine from scratch. The U.S. Copyright Office weighed in with a pre-publication report on generative AI and training, available at copyright.gov, which analyzed how "fair use" applies when AI systems train on copyrighted works. The report's careful position, essentially that some training uses may qualify as fair use while others, particularly those producing substitutable outputs, may not, gives both sides language they can quote and neither side a clean win. Judges reading the new complaints will have that analysis in the file alongside the parties' briefs.
Why Regional Papers, Why Now
The timing and the plaintiff profile both deserve attention. Regional papers have been the industry's most damaged survivors over the last two decades; they have the archives (cheap to digitize, expensive to produce) and the standing (registered copyrights on a large fraction of their back catalogs) without the conglomerate resources that let bigger publishers negotiate licensing deals from strength. Filing a lawsuit is one of the few leverage points available to a mid-size publisher facing a company valued in the hundreds of billions.
For OpenAI, each new plaintiff adds discovery burden, potential liability exposure, and reputational friction with the news ecosystem it needs for a viable answer-engine product. For publishers, the suits create pressure for licensing markets that largely did not exist two years ago: AP deals, individual publisher agreements, and content-exchange frameworks have all emerged in the shadow of litigation. Whether that market scales to include regional papers, or whether it concentrates among a few big licensors, is a question the courtroom will not answer.
What is Not Settled
A dose of honesty about the record here: the specific legal theories, exhibits, and damages claims in the Seattle Times and Newsday filings are not fully described in the available reporting, and I will not pretend to know the complaint's internals. What the techbuzz.ai article establishes is the existence of the suits, the defendants, and the framing of the dispute. The rest, on jurisdiction, on what licenses each publisher previously had (if any), on how the training data was actually sourced, remains to be litigated into the record.
What the new filings do change is the shape of the argument. A dispute between one national paper and two AI companies is a fight. A dispute between one national paper, two regional papers, and a growing roster of publishers across the country is a policy problem wearing a litigation costume. Congress could resolve it with a licensing framework or a statutory exception; the courts will resolve it anyway, one complaint at a time, unless lawmakers decide that this is a question better answered in a statute than in a decade of discovery.
Either way, the Seattle Times and Newsday have made the stakes concrete for a segment of the industry, regional and local news, that tends to get ignored in AI policy debates dominated by national players and national headlines. The model providers will have to answer to them too.
Marcus Chen-Ramirez Senior Technology Correspondent, Buzzrag
More Like This
Claude Marketing Skills Ranked by GitHub Stars (2026)
Which Claude Code marketing skill repos actually earn their stars? We map the top packages—from CRO to paid media—and ask what GitHub popularity really measures.
AI's Inference Crisis: Why Sora Died Burning $15M Daily
OpenAI killed Sora after six months. The reason reveals AI's shift from training races to inference economics—and what breaks next.
Tech Career Decisions: What to Know Before 2026
Marina Wyss breaks down seven tech roles—from software engineering to applied science—through a decision tree based on personality, not just skills.
Data Is Now the Hard Part of Building AI
At a recent YC Paper Club session, three AI researchers made the case that training data—not models or chips—is where the real work of building AI happens now.
AI Training Data: The Legal Vacuum No One Is Filling
AI companies built trillion-dollar products on data they never paid for. Courts are starting to push back—and Congress still hasn't shown up.
Inside Anthropic's Project to Scan Millions of Books for AI
Anthropic's Project Panama destroyed hundreds of thousands of books to train Claude. Here's how AI companies are turning literature into training data.
Claude Mythos: Hype, Leaks, and What Anthropic Said
A Mythos identifier briefly appeared on Anthropic's API, then vanished. Here's what that actually tells us—and what it doesn't—about a public release.
MemPalace Gives AI Coding Agents Long-Term Memory
MemPalace stores your AI coding conversations word-for-word, locally. Here's what it actually does, where it falls short, and who it's built for.
RAG·vector embedding
2026-09-06This article is indexed as a 1536-dimensional vector for semantic retrieval. Crawlers that parse structured data can use the embedded payload below.