Unsealed court filings show a Microsoft scientist called AI training data collection an astonishing theft

Started by Jeffy, Yesterday at 01:37 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: Unsealed court filings show a Microsoft scientist called AI training data collection an astonishing theft   Views(Read 67 times)
Active members in this topic:
Jeffy(1)

Jeffy

The New York Times' ongoing copyright lawsuit against Microsoft and OpenAI has produced a genuinely damaging set of internal statements now that a partly unredacted court brief has become public, and the most striking line in the whole filing comes from inside Microsoft itself. Brent Hecht, the company's director of applied science, described the mass collection of data used to train AI systems as an astonishing theft of unprecedented proportions, going as far as calling it possibly the largest theft of labor in human history, language that reads far more like something from a Times press release than an internal Microsoft document.

Microsoft was not the only company caught making uncomfortably candid internal admissions. OpenAI's head of ChatGPT, Nick Turley, is quoted describing the company's own products as largely substitutive, period, going on to warn that publishers faced what he termed an existential threat from exactly the kind of AI generated answers that increasingly replace a visit to the original news source. Microsoft CEO Satya Nadella himself testified separately that chatbot conversations had substituted for publisher sites, a direct executive level acknowledgment of the same substitution dynamic Turley described from inside OpenAI.

Microsoft's own internal analysis reportedly went further still, documenting a self reinforcing problem the company internally referred to as a doom loop, in which AI generated answers reduce the website traffic that originally supplied the training data those answers depend on in the first place, a genuinely circular problem where the technology's own success gradually starves the ecosystem it was trained on.

The legal stakes riding on these particular admissions are significant. The Times originally filed this lawsuit back in December 2023, alleging unauthorised use of its journalism in training AI systems, and the case has since been folded into broader multidistrict proceedings alongside similar claims from other publishers. The court has already dismissed the contributory infringement and trademark claims specifically, narrowing what remains to direct infringement and the core fair use question, which is exactly the area these newly unsealed internal statements are aimed squarely at. The Times is using Hecht's and Turley's own words to argue that both companies clearly understood their products could function as direct market substitutes for the very material they trained on, a factor that weighs heavily against a fair use defence under the purpose and market effect prongs of that legal test.

Both companies are pushing back on how much weight these statements should actually carry. Microsoft maintains that Hecht's comments represented his own personal perspective rather than any kind of official corporate legal position, while OpenAI has largely stayed quiet on the specific quotes and continues to maintain more broadly that AI training itself qualifies as fair use regardless of what individual employees may have said internally. Whether a judge finds an executive level and a senior scientist level admission this candid more persuasive than that formal legal defence is likely to be one of the more consequential questions in the entire AI copyright litigation landscape currently working through American courts.

Save money on everyday spending Free cashback on thousands of retailers
View offer