← Back to all posts
News

Unsealed New York Times Brief Shows Microsoft Measured a 93% Copilot Click-Through Drop

September 20, 2026 · 03:13 UTC · News
Unsealed New York Times Brief Shows Microsoft Measured a 93% Copilot Click-Through Drop

TL;DR

On September 17, a partly unredacted version of the news publishers' summary-judgment brief in In re OpenAI, Inc., Copyright Infringement Litigation was unsealed in the Southern District of New York. The quotes inside it are brutal. A Microsoft director of applied science called the scraping "an astonishing theft of unprecedented proportions" and possibly "the largest theft of labor in human history." OpenAI's head of ChatGPT told colleagues that publishers faced an "existential threat" because the products are "largely substitutive."

Those lines went everywhere. They are not the part that should change how you build. Sitting in the same filing is a figure produced by Microsoft's own product telemetry: click-through to nytimes.com fell by as much as 93% when Copilot's answer engine handled a query instead of a conventional Bing results page.

Opinions are cheap in discovery. A substitution rate the defendant measured on its own dashboard, before anyone had to defend it, is not.


The 93% Is the Evidence. The Insults Are Decoration.

Fair use in the United States runs on four factors under 17 U.S.C. section 107. The fourth one asks whether the use harms the market for the original work, and in practice it decides a lot of these cases. Everything else in the brief is context. The 93% goes straight at factor four.

What makes it hard to wave away is its provenance. This is not a plaintiff's expert modelling a counterfactual web. It is Microsoft instrumentation on a Microsoft surface, presented internally by a Microsoft scientist on a slide, comparing traffic sent to a publisher domain by two of its own products.

click-through to nytimes.com (bing search = 100) Bing search100 Copilot answers7 source: microsoft internal telemetry cited in the unsealed brief
A drop of up to 93%, measured by the company whose product caused it.

Think of a restaurant that starts handing out free tasting spoons of the dish next door, then checks its own door counter and finds that almost nobody walks next door anymore. The counter was installed for operations, not for a lawsuit. That is exactly why it is awkward to explain away later.

Project Taxi and Project Mango

The brief also puts numbers and internal code names on the content pipeline. According to the unsealed filing, Microsoft and OpenAI exchanged training material through efforts called Project Taxi and Project Mango, and the Project Mango data was assembled into a training set containing copies of at least 160,903 unique works from the plaintiff publishers.

Separately, a Common Crawl derived dataset is described as holding more than two million documents from nytimes.com, and OpenAI mid-training datasets are alleged to contain 91,692 copies of works from the Times, the New York Daily News and the Center for Investigative Reporting.

the content path described in the unsealed brief Common Crawl2M+ nytimes.com docs Project Taxi / Mango160,903 works training sets91,692 copies counts are the publishers' allegations, not findings of fact
Code names and counts the public had never seen attached to each other before.

Two further allegations in the brief matter more than the counts. The publishers say paywalls were circumvented, pointing to an exchange in which OpenAI researcher Nick Ryder described a method to president Greg Brockman, who replied "ah nice." And they say copyright notices were stripped from material before it went into training, which is the territory covered by 17 U.S.C. section 1202 on copyright management information.

What the Executives Actually Said

The headline quotes belong to Brent Hecht, a Microsoft director of applied science, who wrote internally that the copying amounted to "an astonishing theft of unprecedented proportions" and perhaps "the largest theft of labor in human history." A later internal presentation of his describes a doom loop for the web and puts the supply problem plainly: "It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its 'content supply chain.'"

OpenAI's Nick Turley, who runs ChatGPT, is quoted warning that publishers face an "existential threat" and that the products are "largely substitutive." Microsoft CEO Satya Nadella testified in deposition that "anything that is paywalled should be licensed by anyone who wants to use it," and that chatbot conversations had substituted for visits to publisher sites.

Microsoft's response is the one you would write too. Spokesperson Alex Haurek told The Verge that the Hecht documents "reflect one employee's individual perspective, are not a legal analysis, and do not represent the company's views," and that Nadella "spoke to broad principles and changes underway in how people find and consume information," which "should not be confused with conclusions about copyright questions before the Court." OpenAI and Microsoft did not respond to requests for comment from TechCrunch.

Steven Lieberman, counsel for the New York Daily News, framed it the other way: "The evidence revealed here for the first time shows that OpenAI and Microsoft knew that what they were doing was wrong."

Why a Builder Should Care

Three concrete things change on your side of the fence.

1. Your own analytics are discoverable

If you ship a product that summarizes, answers over, or otherwise stands in front of somebody else's content, the substitution dashboard you built for retention is also an exhibit. Nobody at Microsoft instrumented that Copilot comparison in order to lose a lawsuit. They instrumented it because product teams measure things, and then discovery happened.

2. "Transformative" is not a free pass on factor four

A lot of AI product planning still assumes the fair use conversation ends at whether the output looks different from the input. This filing is a reminder that a court also gets to ask whether the thing you shipped ate the market for the thing you trained on, and that a defendant's internal admission on that point is worth more than any expert either side hires.

3. Provenance hygiene is cheap now and expensive later

The copyright-notice allegation is the quiet one. Stripping attribution metadata during a preprocessing step is the kind of decision an engineer makes in an afternoon to normalize a corpus. Keep the provenance fields. Keep the fetch logs. Keep the robots and paywall handling auditable. None of that is expensive until it is.

The Caveats, Straight

This is a plaintiffs' brief. It is argument, written to persuade, quoting selected internal material. Some exhibits remain sealed. Nothing here is a finding of fact, no ruling has issued, the case sits at summary judgment before Judge Sidney H. Stein, and the underlying complaint dates to December 2023. Both companies dispute the characterizations.

What is not in dispute is that the documents exist and that the telemetry was Microsoft's. Whether 93% ends up deciding anything is for the court. Whether it should change how you reason about the market-harm factor is for you, and it should.

Key Takeaways

  • The 93% is the story. Microsoft's own telemetry showed click-through to nytimes.com falling by as much as 93% when Copilot answered instead of Bing, which lands directly on the market-harm factor of fair use.
  • Named pipelines, real counts. The brief describes Project Taxi and Project Mango, a Project Mango training set with at least 160,903 unique publisher works, 2 million-plus nytimes.com documents from a Common Crawl derived set, and 91,692 copies in OpenAI mid-training data.
  • Internal quotes cut against the defense. "Largest theft of labor in human history" from a Microsoft director, "existential threat" and "largely substitutive" from OpenAI's head of ChatGPT, and a CEO deposition saying paywalled material should be licensed.
  • Two allegations beyond volume. Paywall circumvention, complete with a "ah nice" reply from OpenAI's president, and stripping of copyright notices from training material.
  • Your dashboards are exhibits. If your product measures how often it removes the need to click through to a source, assume that measurement is one subpoena away from a filing.
  • It is a brief, not a ruling. Argument from one side, exhibits partly sealed, summary judgment pending before Judge Stein in the Southern District of New York.

Sources: TechCrunch, CourtListener docket, In re OpenAI, Inc., Copyright Infringement Litigation (S.D.N.Y.), The Washington Post, Futurism, TheWrap, 17 U.S.C. section 107

AICopyrightFair UseOpenAIMicrosoftPublishingTraining DataLitigation
CONSOLE
$