← Back to all posts
News

The DOJ Told a Judge AI Training Is Fair Use. Its Example Is Joan Didion.

September 3, 2026 · 02:18 UTC · News
The DOJ Told a Judge AI Training Is Fair Use. Its Example Is Joan Didion.

TL;DR

On September 1 the US Department of Justice filed a 20-page statement of interest in In re OpenAI, Inc. Copyright Infringement Litigation, the consolidated Southern District of New York case in which the New York Times, other publishers and book authors are suing OpenAI and Microsoft. The government's position: copying written works to train a large language model is "extraordinarily transformative" fair use, and rules to the contrary "threaten national security." Per Betanews, it is the first time the federal government has taken a formal side in the AI training suits. The brief calls the "market dilution" theory from Kadrey v. Meta "deeply flawed," dismisses the Copyright Office's 2025 report as "threadbare," and invokes a teenage Joan Didion typing out Hemingway. The Times says Washington sided with "a handful of trillion-dollar AI companies." A statement of interest binds nobody. Summary judgment motions are due Friday.


What was filed

The document is signed by Associate Attorney General Stanley E. Woodward Jr., Assistant Attorney General Brett Shumate of the Civil Division, and Senior Counsel Michael Weisbuch, and filed under 28 U.S.C. § 517, which lets the department "attend to the interests of the United States" in pending suits. It applies to "All Matters" in MDL 25-md-3143 before Judge Sidney Stein, and a footnote extends its arguments to "all parties in this litigation and the related cases, including book authors and publishers."

The framing is national policy first, copyright second. The brief quotes two executive orders, including the January 2025 one on "global AI dominance," and the White House's March 2026 National Policy Framework for AI, then gets to the point: "The United States has a strong interest in this Court rejecting any argument that training LLMs on copyrighted texts violates copyright law."

Per the New York Daily News, Woodward described the brief as historic and said the administration would not let the country be disadvantaged "based on a plainly incorrect understanding of copyright law." The same report and the Times's own coverage, syndicated at GV Wire, note that both sides' summary judgment motions are due Friday. The brief landed three days before the deadline, which is not a coincidence.

Three stages, one of which the government cares about

The brief borrows Judge Stein's own description of how LLMs get built, splits the lifecycle into acquisition, training and output, notes that "each stage may present distinct questions of copyright law," then limits itself to the middle one. That is the move that won in Bartz v. Anthropic, where Judge Alsup found training on books fair use but sourcing them from pirate sites infringing, a distinction that cost Anthropic a $1.5 billion settlement, about $3,000 a book.

three stages, three separate copyright questions (per the brief) Acquisitionhow copies were obtained Trainingcopy to learn patterns Outputwhat users get back Bartz: piracy cost $1.5B DOJ: fair use, full stop judged output by output the brief takes a position only on the middle box
The government's argument is deliberately narrow: training is its own use, and it is fair.

Copyright judges the act, not the object. The same photocopied chapter is fair when a student annotates it and infringing when a shop sells it, and the court asks about each act separately. The brief leans on that "use-by-use" rule, citing Warhol v. Goldsmith and Google v. Oracle, to seal training off from whatever the model later says. On the first factor it borrows Alsup's line that training is "spectacularly" transformative and concludes that "the use of copies to train LLMs is extraordinarily transformative."

Outputs get a carve-out, not a pass. The brief concedes that "certain uses may not be transformative if the LLM reconstructs and disseminates an original copyrighted work," but insists that "whatever legal questions certain output uses might raise, that should not bear on the transformative nature (or any other aspect) of the training use." Footnote 15 previews the remedy fight: "A tiny sliver of anomalous reconstructive outputs would not support a remedy that cuts off or threatens massive liability for LLM output uses generally." Regurgitation, in other words, is a bug to be litigated one output at a time, not a reason to unwind the model.

The fight over factor four

The interesting part of the brief is not that it says fair use. OpenAI has said that since 2023. It is that the government picks a fight with the one federal ruling that gave creators hope. In Kadrey v. Meta (June 25, 2025), Judge Chhabria ruled for Meta because the authors had not built a record, but wrote that "market dilution," the idea that a model trained on your books could flood your market with competing works, "will often cause plaintiffs to decisively win the fourth factor" in future cases. Every publisher complaint since has quoted him.

The DOJ calls that "contrary dicta" that "misapplies copyright principles to LLM training," and says the Kadrey court's fourth-factor analysis is "deeply flawed" because it "improperly collapsed LLM training and LLM outputs into a single continuous use." The government's counter-rule: only outputs that are substantially similar to protected expression can cause cognizable market harm, and "the potential for future outputs that might cause market harm is simply not relevant to evaluating an LLM training use."

fair use factor four: two theories of market harm Kadrey v. Meta, Jun 2025 "market dilution": outputs could flood the market, so developers should pay for training DOJ brief, Sep 2026 "substitutive" harm only: an output must copy protected expression to count against training DOJ on Kadrey: "contrary dicta" :: "deeply flawed" :: "collapsed" two uses
Same statute, opposite readings. The Third Circuit's pending Ross decision may pick one before Stein does.

Then comes the literary exhibit. The brief recounts, citing a 1978 Paris Review interview, that as a teenager Joan Didion "would type out" Hemingway's "stories to learn how the sentences worked." By Kadrey's logic, the government argues, "Didion should have incurred liability to Hemingway every time she published a piece," because the process by which she trained herself and the process by which she produced work "was all one use." It is a good line, and it is the whole dispute in miniature: is a model reading a million articles more like Didion at a typewriter or a photocopier with a distribution deal? The brief asserts the former.

The Copyright Office gets a footnote

In May 2025 the US Copyright Office pre-published Part 3 of its AI report, which said some training would be fair use and some would not, and took the dilution concern seriously. The Register of Copyrights, Shira Perlmutter, was fired the next day and is suing over it. The DOJ brief handles all of that in footnote 17: the Register, "who is currently challenging her removal, appeared to endorse a similar theory in a report," her "understanding does not warrant deference" under Loper Bright, and her "threadbare reasoning ignored all the caselaw."

The March framework told Congress that although the administration "believes that training of AI models on copyrighted material does not violate copyright laws, it acknowledges arguments to the contrary exist and therefore supports allowing the Courts to resolve this issue." Six months later the executive branch walked into the courtroom to help the courts resolve it.

National security, oligopoly, and the Times's own memo

"Rules of law that make it significantly more difficult to develop a robust AI industry in the United States therefore threaten national security and give a competitive advantage to foreign adversaries who are not so encumbered." A licensing requirement, it argues, would mean "only the largest technology companies might have the capital necessary to pay licensing fees," creating "an oligopoly on LLM training due to licensing entry barriers that function primarily as large subsidies for old mainstream media companies."

Two hedges sit underneath. Footnote 13 says the government "takes no position on whether a licensing regime would be financially or logistically feasible," and concedes that publishers "could enter (and have entered) into licensing agreements" for "real-time, pay-walled, proprietary" content whatever the court decides. The brief is not saying content deals are worthless, only that nobody should be forced into one for the training step.

Footnote 19 cites a Futurism report that the Times lets its own writers use LLMs to "conceptualize and edit" articles, then adds that "independent and start-up publications, as well as ordinary people, can too." Citing your opponent's internal AI policy back at them is the legal equivalent of pulling up their old tweets.

The Times's response, via spokesperson Graham James to the Associated Press: "The Administration is siding with a handful of trillion-dollar AI companies at the expense of the countless American creators whose work they stole." And: "AI companies simply need to pay fairly for the content that makes their products possible, as copyright law requires."

Where this lands on the scoreboard

The case law is split and an appeals court is about to weigh in. Judge Bibas in Delaware ruled against fair use in Thomson Reuters v. Ross in February 2025; the Third Circuit heard argument on June 11 and has not ruled as of the latest reports. Bartz and Kadrey went the other way in June 2025. In Germany, the Munich Regional Court held in November 2025 that GPT models memorizing song lyrics infringed and that the EU's text-and-data-mining exception did not cover it. In June, nearly 400 newspapers filed their own SDNY suit against OpenAI and Microsoft.

AI training and copyright: who has said what, and when Feb 2025Ross: not fair use Jun 2025Bartz: training is fair Jun 2025Kadrey: fair, w/ warning Nov 2025Munich: memorization Jun 2026400 papers sue OpenAI Sep 2026DOJ: fair use pending: Third Circuit in Ross (argued Jun 11) :: SDNY briefs due Friday
Two US district rulings for training, one against, one German ruling against memorization, and now the executive branch on the record.

None of this travels. TNW points out that the EU AI Act's Recital 106 requires any provider placing a model on the EU market to comply with EU copyright rules regardless of where the training happened, and EU law has no fair use, only a text-and-data-mining exception with opt-outs. A fair-use win in Manhattan changes nothing in Munich.

What a builder should take from it

  • Training got a friend; acquisition did not. The brief addresses only the training copy. Bartz's line stands: how you got the corpus is a separate question, and the piracy answer cost $1.5 billion. Keep provenance records for anything you fine-tune on.
  • Outputs are still your problem. The government concedes reconstructive outputs "may not be transformative." Memorization tests, output filters and citation behavior are the part of the stack the DOJ did not defend, so that is where plaintiffs will aim.
  • Licensing is repositioned, not dead. Footnote 13 treats deals for real-time and paywalled content as a market that exists whatever the court says. If you sell content to model builders, that is the product the government just called legitimate.
  • The EU does not care what Stein decides. If you ship to EU users, GEMA v. OpenAI and Recital 106 are your operating law. Plan for opt-out compliance regardless of the US outcome.
  • A statement of interest is an opinion with a letterhead. It binds nobody. Watch Friday's summary judgment filings and the Third Circuit's Ross decision, the first appellate ruling on training, which will outrank this brief the day it lands.

Key Takeaways

  • The DOJ filed a 20-page statement of interest on September 1 in MDL 25-md-3143 (SDNY, Judge Stein), signed by Associate AG Stanley Woodward Jr., AAG Brett Shumate and Senior Counsel Michael Weisbuch.
  • Its position: copying works to train an LLM is "extraordinarily transformative" fair use. Acquisition and outputs are separate questions it does not address.
  • It calls Kadrey v. Meta's "market dilution" theory "contrary dicta" and "deeply flawed," and says the Copyright Office's 2025 report "does not warrant deference."
  • Policy arguments: licensing mandates "threaten national security" and would hand "an oligopoly on LLM training" to the biggest companies. The brief takes no position on whether licensing is feasible.
  • The brief is persuasive, not binding. Summary judgment motions are due Friday, the Third Circuit's Ross ruling is pending, and EU law (Recital 106, GEMA v. OpenAI) is unaffected.

Sources: DOJ Statement of Interest, 25-md-3143 (S.D.N.Y. Sept. 1, 2026) (PDF), AP: Trump administration backs OpenAI, NY Daily News: DOJ weighs in on OpenAI case, GV Wire (NYT): Justice Dept. sides with OpenAI, Betanews, TheWrap, TNW, White House AI framework (March 20, 2026), Copyright Office Part 3 report (May 2025), Skadden on Bartz and Kadrey, LawSites: Third Circuit argument in Ross, Courthouse News: 400 newspapers sue, NPR on Perlmutter's firing, Paris Review: Joan Didion interview

AICopyrightFair UseOpenAINew York TimesDOJCreator RightsPolicy
CONSOLE
$