AI companies have long relied on online content. Now they are looking for something high-quality training data, they may be turning to books — and tearing them apart
AI companies are reportedly seeking higher-quality training data from printed books, raising stronger copyright governance questions.
GS2GS3The Indian ExpressGS3AI training dataCopyright and licensingEthics of data use
What happened: AI companies reportedly turn to book-derived training data
AI companies have historically relied on online content for training machine-learning systems. The reported development is that some AI companies are seeking higher-quality training data and are reportedly turning to printed books as a text source for training datasets.
Background and earlier position: web-based training-data dependence
Web-based text has been a practical baseline for training datasets because it is widely accessible and easier to collect at scale. Web text quality also varies, which can motivate searches for more consistent or higher-quality sources.
What changed now: books enter the training pipeline