What ethical and legal safeguards should govern how AI companies source training data? Suggest a regulatory framework applicable to India.
Training large AI models depends on vast human-authored text. Reports that leading firms bought rare books, sliced off their spines, scanned the pages and discarded the remains [1] show that data acquisition has outpaced both ethics and copyright law, making safeguards over provenance, consent and accountability urgent.
Ethical safeguards
- Lawful provenance: datasets built only from legitimately acquired material, never from pirate repositories; internal sourcing projects should not stay secret until litigation exposes them [1].
- Transparency and accountability: disclosure of corpus composition, aligning with NITI Aayog's Responsible AI principles of transparency, accountability and equality [2].
- Creator fairness: authors and publishers compensated, since their work is the input to a commercial product.
- Cultural stewardship: digitisation must be non-destructive; irreplaceable editions are public heritage, not disposable feedstock.
Legal safeguards
- Copyright clarity: India's Copyright Act, 1957 offers only a closed list of fair dealing exceptions under Section 52, with no text-and-data-mining provision — leaving AI training legally untested [3].
- Personal data discipline: web-scraped corpora carry personal data, attracting the DPDP Act, 2023 duties of consent, purpose limitation and notice [4].
- Remedies: statutory damages and injunctive relief, so infringement is not merely a cost of doing business.
A regulatory framework for India
- A statutory TDM exception, conditional on lawful acquisition, non-destructive copying and opt-out for rights-holders.
- A dataset provenance register with mandatory disclosure and third-party audit, operationalised through the AI Safety Institute and AI Governance Group created under the India AI Governance Guidelines (2025) [5].
- Collective licensing via copyright societies, creating a royalty pool for authors.
- A national digital deposit so any scanned rare work enters public archives, and the physical copy survives.
India can convert this global controversy into first-mover advantage: a techno-legal regime that keeps data lawful, creators paid and heritage intact will make #AIforAll credible, letting innovation and cultural preservation advance together rather than at each other's cost.
Sources
- 1How AI firms are destroying physical books to train their models — The Hindu (26 August 2026)destructive scanning of purchased rare books; secrecy of AI firms' book-sourcing programmes
- 2Responsible AI #AIForAll: Approach Document for India, Part 1 — Principles for Responsible AI, NITI Aayogtransparency, accountability and equality principles; #AIforAll framing
- 3The Copyright Act, 1957 — India CodeSection 52 fair dealing as an enumerated, closed list of exceptions
- 4The Digital Personal Data Protection Act, 2023 (No. 22 of 2023) — MeitYconsent, notice and purpose limitation obligations on data fiduciaries
- 5MeitY unveils India AI Governance Guidelines under IndiaAI Mission — PIB (5 November 2025)AI Governance Group, Technology & Policy Expert Committee and AI Safety Institute; techno-legal approach