Skip to main content
AI

AI scraping is theft, admits a Microsoft exec

A senior Microsoft exec called the AI industry’s mass scraping “an astonishing theft”—and it could lead to more lawsuits.

3 min read

TOPICS: AI / AI Governance / Copyright & Intellectual Property

TL;DR: It’s the “largest theft of labor in human history.” That’s what a senior Microsoft exec said about AI scraping in newly unsealed filings from The New York Times’ copyright lawsuit. In the same filings, OpenAI’s head of ChatGPT warned that the company’s product was an “existential threat” to publishers. This could all lead to more lawsuits—just as courts have been siding with the AI companies.

What happened: Millions of people would consider AI models “hoovering up” their work as an “astonishing theft of unprecedented proportions,” Microsoft Director of Applied Science Brent Hecht wrote in an internal memo in 2023. This came to light after the NYT sued OpenAI and Microsoft for copyright infringement later that year.

But what does “unprecedented proportions” look like? Per the lawsuit, more than 91,692 copies of work from the NYT and additional plaintiffs are included in OpenAI’s mid-training datasets (curated sets of high-value data) alone. Project Mango, a joint data-sharing effort between Microsoft and OpenAI, allegedly produced a training set with at least 160,903 unique works from news publishers.

Paywall? No problem: OpenAI employees allegedly devised plans to hop paywalls to get training data. Per the filing, when an OpenAI researcher flagged a hack to get around the Times’ paywall, Greg Brockman, co-founder and president of OpenAI, replied: “ah nice.” Microsoft CEO Satya Nadella testified earlier this year that if he’d been made aware OpenAI was scraping behind paywalls, he’d have forced the company to retrain its models: “Anything that is paywalled should be licensed by anyone who wants to use it.”

Tech news that makes sense of your fast-moving world.

Tech Brew breaks down the biggest tech news, emerging innovations, workplace tools, and cultural trends so you can understand what's new and why it matters.

By subscribing, you accept our Terms & Privacy Policy.

Fair use: While this could all lead to more lawsuits, AI companies have been winning similar cases on fair use grounds (a law that allows the unlicensed use of copyrighted material under certain circumstances). In a case last year, a judge sided with Anthropic over its digitization and destruction of legally purchased print books to train its models, though the ruling was split—it later paid $1.5 billion in a settlement to authors over millions of pirated books.

OpenAI’s warning of an “existential threat” could also undermine its defense: Head of ChatGPT Nick Turley described the company’s products as “largely substitutive” and said they’d get more so as they improved (in other words, why go to the source when the chatbot told you what you needed?)—cutting against the law’s requirement that the use doesn’t substitute the original work.

Bottom line: Copyright lawsuits against AI companies are piling up—and more are likely on the way. Now, authors and news publishers can use the companies’ own words to make their case. —LC

About the author

Lindsey Choo

Tech Brew

Tech Brew breaks down the biggest tech news, emerging innovations, workplace tools, and cultural trends so you can understand what's new and why it matters.

By subscribing, you accept our Terms & Privacy Policy.