·The Hindu·15 marks·250–350 words

"AI training data acquisition is testing the limits of existing copyright frameworks." Discuss with reference to recent global controversies, and examine implications for India's IPR regime.

In this answer
  1. Global controversies exposing the strain
  2. Implications for India's IPR regime

Copyright law was built for human authors and discrete acts of copying; large language models instead ingest entire corpora. Recent disclosures that AI firms destructively scanned physical books — cutting spines and discarding them after digitisation [2] — show acquisition practices outpacing statutory design.

Global controversies exposing the strain

  • Piracy-sourced datasets: in Bartz v. Anthropic PBC (N.D. California), authors alleged training on books from shadow libraries LibGen and PiLiMi; the court held digitising legally purchased books to be transformative fair use, but pirated sourcing unprotected [1].
  • Record liability: the resulting $1.5 billion settlement (approved 2026) is the largest copyright recovery in US history — pricing litigation risk into AI data sourcing [1].
  • Corporate opacity: Anthropic's internal book-scanning programme and Amazon's warehouse scanning surfaced only through litigation and investigative reporting, not disclosure [2].
  • Heritage cost: rare and out-of-print editions are destroyed post-scan, an irreversible cultural loss no copyright remedy repairs [2].

Implications for India's IPR regime

  • Fair dealing is narrower than fair use: Section 52, Copyright Act 1957 lists closed-ended purposes, unlike the open US standard — creating uncertainty over whether training is "research" [3].
  • Judicial testing has begun: in ANI Media v. OpenAI, the Delhi High Court declined an interim injunction, treating LLM training as prima facie fair dealing pending final hearing [4].
  • Legislative gap: the Parliamentary Standing Committee on Commerce (161st Report, 2021) flagged the absence of a framework for AI-related works and urged review of the IPR statutes [3].
  • Data-governance overlap: the Digital Personal Data Protection Act, 2023 governs personal data but not copyrighted training corpora, leaving text-and-data-mining unregulated [5].

India should legislate a calibrated text-and-data-mining exception with mandatory dataset disclosure, lawful-acquisition conditions and creator remuneration, aligning innovation under the IndiaAI Mission with authors' Article 300A property interests. A transparent, licensed data market — not destruction in secrecy — best serves both creators and India's AI ambitions.

Sources

  1. 1Bartz v. Anthropic PBC, No. 3:24-cv-05417 (N.D. Cal.) — docket and ordersfair-use ruling on legally purchased books, piracy exclusion, $1.5 billion settlement
  2. 2How AI firms are destroying physical books to train their models — The Hindu explainer, 26 August 2026destructive spine-cutting and scanning, corporate non-disclosure, heritage loss
  3. 3Standing Committee on Commerce, 161st Report: Review of the Intellectual Property Rights Regime in India (2021) — PRS summarystatutory gaps on AI and copyright review
  4. 4ANI Media Pvt. Ltd. v. OpenAI Inc. — Delhi High Courtinterim refusal of injunction; training treated as prima facie fair dealing
  5. 5The Digital Personal Data Protection Act, 2023 (Act 22 of 2023), MeitYpersonal-data scope, no coverage of copyrighted training corpora

More from this note