AI companies destroy physical books – let's scan rare books before it's too late

579 points · 861 comments on HN · read original →

Points and comments are a snapshot, not live.

AI companies are buying, scanning, and destroying physical books to train models, locking knowledge on private servers.

Anthropic's 'Project Panama' spent tens of millions purchasing millions of paper books for scanning, then destroyed them to train Claude LLM, as exposed in a $1.5 billion copyright settlement. The practice prevents competitor access, avoids legal risks, and costs less than lossless scanning. Anna's Archive calls for volunteers worldwide to scan and upload rare materials before publishers or AI companies eliminate public access, offering recognition and scanning fee reimbursement for large-scale contributions.

What commenters are saying

Many commenters blame copyright law and publishers for locking up books, forcing AI companies into destructive scanning. A split emerges: some argue AI firms could avoid destroying books if they paid for e-books or licensed content, while others note that destroyed books are often commercially dead inventory, not rare. Legal nuance is cited: Judge Alsup ruled that scanned copies are fair use only if originals are destroyed, leaving a single digital copy on company servers. Skeptics question the rarity of targeted books.