What happened

AI firms are exploring digitised books as a source of text for AI training datasets. The reported rationale combines a need for larger and more reliable datasets with increasing legal and regulatory constraints affecting training-data sourcing.

Background and earlier position

AI firms have commonly used web-accessible digital material as a major input source for AI training. Under increasing scrutiny, AI training pipelines are described as seeking clearer rights status for content used in training, including through permissions and licensing, digitisation arrangements, or controlled access mechanisms.

What changed now

The reported shift is increased interest in books as training inputs and a movement of dataset sourcing toward physical books that are converted into text datasets through digitisation.