Investigate what happens when commercial LLMs crawl, embed, and synthesize proprietary scholarly research. Analyze the legal divide between institutional licenses, US Title 17 § 107 Fair Use, EU TDM opt-outs, and the $35-$45/article access barrier.
Under Authors Guild v. Google (2015), full-text scanning for non-expressive index creation was deemed transformative. However, generative LLMs that synthesize verbatim text or act as commercial substitutes for paywalled publisher research encounter severe friction under the 4th factor (Market Harm to Publisher Licensing Markets, per NYT v. OpenAI and Andersen v. Stability AI).
Article 3: Unconditional, mandatory TDM exception for research organizations and cultural heritage institutions (non-commercial).
Article 4: Permits commercial TDM only if the rightholder has not expressly reserved rights (the "opt-out" mechanism via machine-readable standards like robots.txt or TDM reservation protocols).
Limits text and data analysis strictly to non-commercial research with lawful access. Proposed broader commercial AI exceptions were halted in 2023 following vehement opposition from academic publishers and creative industry lobbies.
Independent of whether text ingestion constitutes Fair Use under § 107, bypassing technical protection measures (TPMs)—such as Cloudflare challenge screens, authentication tokens, IP bans, or publisher paywall shields—triggers strict liability under § 1201(a)(1). Fair Use is generally not an affirmative defense to a § 1201 circumvention violation.
Following hiQ Labs v. LinkedIn (9th Cir. 2022), scraping publicly viewable data without authentication does not violate the CFAA. However, scraping authenticated content or violating explicit publisher click-through subscription contracts creates immediate breach of contract liability.
Google scanned 20M+ library books to deliver search snippets. Held: highly transformative, non-expressive use that did not substitute for original books because only snippets were displayed to the public.
Key issue: Generative models regurgitating substantial portions of paywalled journalism without compensation, destroying licensing channels. Publishers argue LLM synthesis is expressive substitution, not non-expressive search indexing.
Statutory default judgments establishing tens of millions in willful statutory damages under US copyright law for bulk unauthorized dissemination of 80M+ research papers.