·The Hindu·15 marks·250–350 words

Examine the copyright and ethical challenges posed by the use of physical books for training Large Language Models. What policy safeguards are needed?

In this answer
  1. Copyright challenges
  2. Ethical challenges
  3. Policy safeguards needed

Large Language Models require vast text corpora, and firms have begun buying rare and used print books through front entities, slicing their spines, scanning the pages and pulping the originals — acquiring training data while sidestepping e-book licensing [1]. This "buy-scan-destroy" model exposes deep gaps in copyright law and raises distinct ethical concerns.

Copyright challenges

  • Fair use ambiguity: a US court held in Bartz v. Anthropic (June 2025) that using purchased books to train LLMs was "quintessentially transformative" fair use, while retaining pirated downloads in a general library was not [2]. The split verdict effectively rewards physical acquisition-and-destruction.
  • Indian law is silent on AI: the Copyright Act, 1957 protects original literary works and permits only enumerated "fair dealing" uses; machine ingestion for model training fits no existing exception [3].
  • No remuneration for authors: the first-sale doctrine ends the author's claim at purchase, so writers and publishers gain nothing when one copy trains a commercial model used by millions.
  • Opaque provenance: shell-named buyers such as the "Red Sparrow" and "Blue Finch" projects make datasets untraceable, defeating enforcement [1].

Ethical challenges

  • Destruction of cultural artefacts: pulping rare and out-of-print editions permanently removes physical copies from the public domain of reading.
  • Consent and attribution: authors are neither informed nor credited, reducing creative labour to raw material.
  • Concentration of power: only well-capitalised firms can buy libraries at scale, widening the digital divide in AI capability.

Policy safeguards needed

  • A statutory text-and-data-mining exception with mandatory licensing and royalty-sharing for rights-holders.
  • Training-data disclosure norms, consistent with the algorithmic transparency and data-management pillars of the India AI Governance Guidelines [4].
  • Deposit-before-destruction: mandatory transfer of scanned rare titles to the National Library or AIKosh-type public repositories [4].

Copyright must evolve from a reproduction-centric regime to one governing machine learning. Aligning the Copyright Act with India's AI governance framework can secure both innovation and the creator's constitutional stake in the fruits of their labour.

Sources

  1. 1"The many ways to destroy a book", The Hindu (10 September 2026)buy-scan-pulp practice, front-entity book buyers
  2. 2Bartz v. Anthropic PBC, N.D. Cal. (docket, June 2025 fair-use order)training held transformative fair use; pirated central library not
  3. 3The Copyright Act, 1957, India Codeprotection of literary works, limited fair-dealing exceptions
  4. 4India AI Governance Guidelines, MeitY/IndiaAI (PIB)data management, algorithmic transparency, AIKosh repositories

More from this note