A newly unredacted court filing in The New York Times' three-year-old copyright lawsuit against OpenAI and Microsoft contains internal quotes that are, on their face, remarkable.

Before any of them: this is one side's brief in active litigation. TechCrunch, which reported it, says so directly — "much of the new information comes from The Times' own brief, not the underlying exhibits, which remain sealed," and the quotes are "presented without their original context." OpenAI and Microsoft did not return requests for comment.

Read on with that held firmly in place. It doesn't make the material unimportant. It makes it evidence rather than verdict.

The Quote

Per the filing, in a January 2023 internal memo, Brent Hecht — Microsoft's director of Applied Science — described AI scraping as:

"an astonishing theft of unprecedented proportions"

and

"the largest theft of labor in human history."

Worth being precise about what that is and isn't. It's a named senior scientist's private assessment, quoted by opposing counsel. It is not a Microsoft corporate position, and Microsoft's public legal argument has been the opposite: that training on copyrighted material is fair use.

Those can both be real. Large companies contain people who disagree with the company's lawyers. That's normally invisible — until discovery.

The Part That Probably Matters More

The memorable quote will travel. The business analysis is the material with legal weight.

Per TechCrunch, "Microsoft's own data shows its Copilot 'answer engine' caused click-through rates for The New York Times' domain to drop as much as 93% compared to traditional Bing search." A January 2024 internal presentation by Hecht called this a "doom loop" that would "hurt the performance of our models and the entire web at the same time."

And this, from the Microsoft document as quoted in the filing:

"It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its 'content supply chain.'"

That sentence is doing something specific. It's not a moral complaint. It's a supply-chain observation: the product degrades the thing it's built from. If publishers can't fund reporting, there's less reporting to train on, and the models get worse. A doom loop.

Other quoted material points the same way. OpenAI's head of ChatGPT, Nick Turley, reportedly wrote that publishers face an "existential threat" from products that are "largely substitutive" and "will get more and more substitutive as they get better." A Microsoft document reportedly warned of a "real risk" that generative AI could "significantly disrupt the employment of the very people who generated the data on which the foundation model was trained."

Why Lawyers Care About That Specific Word

"Substitutive" isn't a casual adjective here. It's close to a legal term of art.

Fair use is a four-factor test, and one factor is the effect on the market for the original work. A use that transforms something — commentary, parody, criticism — sits comfortably. A use that substitutes for it, so people consume the copy instead of the original, sits badly.

So internal documents describing your own product as "largely substitutive," alongside data showing a 93% collapse in click-throughs, land directly on the factor the defense most needs.

That's the strategic logic of unsealing this material now. Whether it works is a different question — and TechCrunch notes that judges "have been largely favorable to AI companies' arguments that training constitutes 'fair use,'" with the Trump administration filing in support of OpenAI's position earlier this month.

The Scale

The filing also puts numbers on the copying. Per TechCrunch, OpenAI's mid-training datasets alone contain more than 91,692 copies of works published by the NYT, the Daily News, and the Center for Investigative Reporting. A Common Crawl-derived dataset reportedly included more than 2 million documents from nytimes.com. A separate dataset assembled under something called Project Mango allegedly contains at least 160,903 unique works from the news publishers.

Two further allegations in the filing, which are about conduct rather than volume:

That OpenAI staff discussed getting around paywalls — the brief reportedly quotes researcher Nick Ryder telling Greg Brockman about a "hack to get around nytimes paywall," with Brockman replying "ah nice."

And that copyright notices were deliberately stripped from training data, because researchers "wouldn't want model outputting" "copyright notices" to users.

If the underlying exhibits support them as characterized, those are the details that look worst — not the volume, but the awareness.

How to Read a Story Like This

(This section is context, not the reporting.)

Here's the media-literacy point, and it's the reason this story is worth your time even if you don't care about copyright law.

A legal brief is an argument, not a record. Its job is to assemble the most damaging accurate things and arrange them persuasively. Every quote here was chosen by people paid to win. Selection is not fabrication — these appear to be real documents — but a memo can read very differently with its surrounding paragraphs attached.

Three habits worth having when a story like this crosses your feed:

Ask who is quoting whom. "Microsoft said X" and "The New York Times' lawyers say a Microsoft document said X" are different sentences. Most coverage will collapse them.

Ask what's still sealed. Here, the exhibits are. That's the check that hasn't happened yet.

Ask whether the other side responded. They didn't, in this case — which is itself unremarkable during litigation, and is not evidence of anything.

None of that is a reason to dismiss it. If the exhibits bear out the brief's characterization, these are exactly the kind of internal documents that shape a landmark case. It's a reason to hold it as strong allegation rather than established fact, which is a distinction worth modelling out loud for anyone who watches news with you.

For Parents and Teachers

There's a question underneath this case that reaches your kitchen table, and it isn't legal.

The material these models learned from was made by people — reporters, photographers, researchers — who were paid because someone visited a page. If the answer arrives without the visit, the funding thins, and eventually so does the material.

Microsoft's own quoted phrase for this is the honest one: a product that "threatens the economic foundations of its essential suppliers."

You don't need a position on the lawsuit to draw the practical line for a teenager: when an answer matters, go to the source. Not only because the source is usually better, but because clicking through is the mechanism that funds the thing being summarized. That's a small habit and it's also, in aggregate, the whole argument in this case.

Source: "Microsoft exec called AI scraping 'the largest theft of labor in human history,' new unredacted filings reveal," TechCrunch, September 17, 2026: https://techcrunch.com/2026/09/17/microsoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history-new-unredacted-filings-reveal/

All quotes above are as reported by TechCrunch from The New York Times' court filing. The underlying exhibits remain sealed. OpenAI and Microsoft did not return requests for comment. No claim here has been ruled on.

The sections marked as context — how to read a litigation story, and the notes for parents and teachers — are ours.

Want practical AI guidance for parents and educators every week? Subscribe: https://www.aibyage.com/?modal=signup&utm_source=beehiiv&utm_medium=newsletter