Court documents unsealed this week reveal that executives at Microsoft and OpenAI expressed deep concerns internally about the scale of web scraping used to train AI models like ChatGPT, with one Microsoft director calling it the “largest theft of labor in human history.”
The comments, part of a lawsuit filed by The New York Times against OpenAI and Microsoft in 2023, were included in court filings but only released in snippets without broader context. Dr. Brent Hecht, Microsoft’s director of Applied Science, warned that the scraping practices could create a “doom loop,” while OpenAI exec Nick Turley said it represented an “existential threat to publishers.”
According to the unredacted materials, both companies bypassed paywalls and erased copyright notices from training data. Internal discussions acknowledged that AI systems were “largely substitutive” to journalism, with an OpenAI software engineer telling colleagues in 2023 that “no matter how prominently we show the links, users won’t click.”
Microsoft’s position is set out in its court filings, which explain why these transformative uses are consistent with copyright law and why Copilot is not a substitute for publishers’ journalism, according to Microsoft spokesperson Alex Haurek.
Read Also: Higher USB drive costs reflect real trade-offs in speed
Lawsuit Could Shape Future of AI Training
The New York Times lawsuit, launched alongside five other writers, argues that big tech companies have broken copyright law by scraping millions of their stories off the internet and using the text without approval or compensation to train advanced large language model (LLM) systems. While similar cases have so far favored AI companies, judges have emphasized that the legal framework around AI use remains unsettled.
These revelations come as courts grapple with applying fair use doctrines, originally designed for purposes like parody or commentary, to AI training. The outcome of this case is being watched closely as a potential benchmark for how future disputes will be resolved.
An internal 2023 Microsoft document warned that “millions of people around the world will soon consider large models ‘hoovering up’ all their work to be an astonishing theft of unprecedented proportions.” Dr. Hecht also said that large AI models “are a product that destroys its supply chain.”
Despite public statements defending their practices, the internal communications suggest awareness of significant ethical and legal risks tied to their data collection methods.
