What happened
AI firms are exploring digitised books as a source of text for AI training datasets. The reported rationale combines a need for larger and more reliable datasets with increasing legal and regulatory constraints affecting training-data sourcing.
Background and earlier position
AI firms have commonly used web-accessible digital material as a major input source for AI training. Under increasing scrutiny, AI training pipelines are described as seeking clearer rights status for content used in training, including through permissions and licensing, digitisation arrangements, or controlled access mechanisms.
What changed now
The reported shift is increased interest in books as training inputs and a movement of dataset sourcing toward physical books that are converted into text datasets through digitisation.
Related current affairs
- AI companies have long relied on online content. Now they are looking for something high-quality training data, they may be turning to books — and tearing them apart
- Cut, copy: Why books are being destroyed to feed AI
- AI companies and online content digitisation (right-column short box)
- What Indian coders say: Impact of book digitisation and destruction (box within feature)
- Government sets up Centres of Excellence in AI for domains like agriculture, healthcare, sustainable cities and education
- Why is China’s AI push gaining ground?
