- Anthropic purchased millions of physical books, removed their spines, scanned every page, and recycled the remaining parts to create training data for its AI models.
- The internal project, known as “Project Panama,” aimed to “destructively scan all the books in the world,” with the company spending tens of millions of dollars on this endeavor.
- The revelations about Anthropic’s book scanning project highlight the fierce competition among AI companies for high-quality training data, particularly from copyrighted books, as legal questions around AI training data continue to be debated in US courts.
Newly unsealed court filings show that artificial intelligence (AI) startup Anthropic bought millions of physical books, removed their spines, scanned every page, and recycled what was left as part of an internal effort to create training data for its Claude AI models
The project internally called “Project Panama” was described in company documents as an effort to “destructively scan all the books in the world.”
According to the filings, Anthropic spent tens of millions of dollars buying books from used-book sellers before dismantling and digitizing them. Internal planning materials also reportedly noted, “We don’t want it to be known that we are working on this.”
Advertisement
The disclosures, which surfaced in a copyright lawsuit, underscore the intense competition among AI companies to secure high-quality training data
Anthropic reportedly considered books especially valuable because they can help models learn to write better than training data pulled from the open internet
The revelations come amid increasing legal scrutiny over whether and how copyrighted books can be used to train AI systems, with key questions about the boundaries of AI training data still being decided in US courts

