What happened: online-content sourcing for AI training and digitised materials from publishers/libraries

AI companies build training datasets using content available on the internet. Reported controversies focus on whether a meaningful portion of the dataset corpus is drawn from digitised materials originally produced by publishers and libraries. Rights holders and users raise concerns that dataset building can proceed even when permission and licensing terms are unclear, especially when the dataset contains copies of content that may be protected under copyright or governed by contracts.

Background and earlier position: dataset building from online sources fuels a recurring rights-control debate

AI training data pipelines often combine multiple sources of text and other digital content available online. When publishers and libraries digitise books, articles, or other works, digitised copies can enter online collections. The continuing ethical and regulatory tension is that AI training may treat online availability as implied permission, while rights holders argue that permission must be explicit, licensed, and aligned with the intended use (training, downstream outputs, and reuse).

A second recurring issue is takedown. Even when an AI company claims compliance through removals after complaints, critics argue that takedown mechanisms are uncertain, vary across platforms and datasets, and may not fully address risks created during the time content was included in training data. As a result, the controversy persists around how datasets are built, what rights govern those datasets, and what happens when permissions are disputed.

What changed now: licensing uncertainty and takedown risk stay central to digitisation controversy