⚖️ AI work theft

Unsealed Admissions in The New York Times v. OpenAI & Microsoft

Newly unsealed court filings in the ongoing copyright lawsuit brought by The New York Times against OpenAI and Microsoft have exposed internal communications that threaten to undermine the defense's core legal arguments. The unredacted brief quotes high-ranking personnel, including Microsoft’s Director of Applied Science, Brent Hecht, who privately characterized mass AI data scraping as "an astonishing theft of unprecedented proportions" and potentially "the largest theft of labor in human history". Internal documents further reveal that OpenAI researchers devised technical workarounds to bypass paywalls, systematically stripped Copyright Management Information (CMI) from training data to avoid echoing notices in outputs, and acknowledged that chatbots act as direct substitutes for publisher websites. Crucially, Microsoft’s internal metrics indicated that its Copilot feature reduced click-through traffic to The New York Times by up to 93%, leading internal researchers to warn that generative AI risked creating a "doom loop" by destroying the supply chain of original content creators upon which foundation models rely.

Judicial Impact on Fair Use and Corporate Intent

For startup founders building, fine-tuning, or deploying artificial intelligence models, these revelations fundamentally alter the legal calculus around the "fair use" defense under Section 107 of the Copyright Act. Fair use balances four statutory factors, with the transformation of the work and its effect on the potential market for the original being paramount. The unsealed evidence directly targets both elements by showing that executives internally viewed their AI systems not as transformative tools, but as substitutive products that actively cannibalized traffic and revenue from content creators. Furthermore, evidence demonstrating intentional paywall circumvention and deliberate removal of copyright notices creates significant exposure under Section 1202 of the Digital Millennium Copyright Act (DMCA), which carries statutory damages independent of traditional copyright infringement. These admissions severely weaken the argument that web scraping is inherently protected, signaling that courts may soon draw sharper boundaries between permissible public indexing and unauthorized commercial model training.

This evidentiary escalation indicates that early-stage startups can no longer rely on unvetted, mass web scraping or open datasets like Common Crawl without assuming substantial legal liability. Founders building AI applications must conduct immediate data provenance audits to confirm that their training, fine-tuning, and Retrieval-Augmented Generation (RAG) pipelines do not ingest paywalled content, bypass technical access controls, or scrub copyright metadata. When acquiring specialized data, startups should secure explicit, written commercial licensing agreements that clearly grant model-training rights rather than assuming fair use will cover the ingestion. Additionally, because foundational model providers face mounting legal risks that could lead to forced retraining or model deprecation, technical teams should architect their software stacks to be model-agnostic, allowing seamless migration between API endpoints. Finally, founders should update customer contracts and investor disclosures to detail data sourcing practices, protecting the business from downstream indemnity claims should upstream data sources face judicial challenge.

In addition to our newsletter we offer 60+ free legal templates for companies in the UK, Canada and the US. These include employment contracts, investment agreements and more