AI training and rare books are at the center of a growing debate over how historical works should be preserved and used. In an online discussion, one Reddit user argued that artificial intelligence companies may be putting valuable details and historical insights at risk by purchasing rare books without making their contents publicly available.
The poster claimed that AI companies often “never share the contents” of scanned books because they do not want competitors to use the same material to train their own AI models. Critics dismissed the argument as “AI-hating propaganda,” but other participants raised similar concerns about the impact of private AI training on access to cultural heritage.
Some commenters said that training AI on rare and forgotten books could make otherwise inaccessible knowledge easier to find online. Others countered that users might only gain access to “distorted, censored, paywalled snippets of ideas” rather than the complete works. One participant described the idea of an AI “literally eating books to become more powerful” as inherently dystopian, even for people who support the technology.
The BBC reports that some booksellers believe not every printed work needs to be preserved indefinitely. However, Scottish bookseller Derek Walker said AI companies should distinguish between obscure books that have surviving copies and genuinely rare historical works that may be “the only known surviving editions.”
“It would be an even bigger problem if something like this that had been around for so long was bought for destruction,” Walker told the BBC.
The controversy highlights the need for greater transparency when AI companies acquire books in bulk for data collection or model training. Buyers who protect their identities and keep training datasets confidential may make it difficult for booksellers, libraries, historians, and the public to understand which works are being digitized, preserved, or potentially destroyed.
Clearer disclosure could help booksellers identify culturally significant titles before they are sold. Without stronger safeguards, the reputational and historical risks of destroying a rare first edition—or removing an irreplaceable work from public access—could outweigh the potential benefits of using it to advance artificial intelligence.
Source: arstechnica.com


