What happened: reports of book destruction to supply AI training text
A news explained feature describes a growing practice in which libraries, publishers, and digitisation pipelines reportedly destroy physical books—or destroy parts of them—to create large amounts of machine-readable text for AI systems. The reported rationale is operational speed and the scale requirement: AI training needs very large text corpora that must be converted into formats computers can process.
The feature links reported destruction to a tension between: - access to information, including access for AI training and related uses; - copyright and licensing constraints that govern what text can be copied or used; - conservation risk, where digitisation leads to disposal of physical cultural material.
Background and earlier position: two digitisation pathways
The feature explains that digitisation for AI-ready text typically follows two pathways:
First pathway (scan physical books): digitisation pipelines scan physical books to convert printed text into machine-readable form. This can require handling the physical item, and the feature describes cases where portions or whole books may be removed from shelves after conversion.
Related current affairs
- AI companies have long relied on online content. Now, though, it may be turning to books.
- What Indian coders say: Impact of book digitisation and destruction (box within feature)
- AI companies have long relied on online content. Now they are looking for something high-quality training data, they may be turning to books — and tearing them apart
