Newly unsealed court documents from the New York Times' copyright lawsuit against OpenAI and Microsoft, released 17 September, contain internal communications that materially undermine both companies' fair-use defence. Microsoft's Director of Applied Science Brent Hecht wrote in an internal message in January 2023 that AI training on web content constituted "an astonishing theft of unprecedented proportions — the largest theft of labor in human history." In the same document set, Microsoft's own research found that its Copilot answer engine reduced clicks to Times content by as much as 93 per cent — a figure that directly challenges the "transformation, not substitution" argument that forms the core of the fair-use defence. OpenAI leadership separately acknowledged internally that the technology was "largely substitutive" of publisher content and posed an "existential threat" to publishers.
The 93 per cent click-reduction figure is the highest documented substitution rate for any major AI system against any single publisher — and it comes from Microsoft's own measurement, not the plaintiff's expert estimates. Internal acknowledgement that the technology substitutes for rather than complements the content it was trained on is exactly the evidentiary standard that copyright plaintiffs need to overcome a fair-use argument. The documents also reveal deliberate paywall circumvention: Microsoft and OpenAI systematically accessed subscriber-only content during training, not only freely available material. Combined, these three elements — internal acknowledgement of substitution, deliberate circumvention of access restrictions, and proprietary measurement of traffic destruction — describe a state of knowledge that is difficult to reconcile with a good-faith fair-use position.
For GEO practitioners and content marketers: this disclosure reshapes the AI licensing negotiation environment. Any brand, publisher, or media company that has delayed AI licensing discussions is now operating with stronger precedent behind them. The 93 per cent click-reduction figure, sourced from the defendant's own files, is now a negotiating anchor for every future discussion about what AI answer engines are worth to the publishers whose content they were built on.