In early 2024, Anthropic quietly initiated an internal operation dubbed Project Panama, described in planning documents as an effort to “destructively scan all the books in the world.” The company spent tens of millions of dollars acquiring physical books, removing their spines with industrial machinery, scanning the pages, and recycling the remnants—all to feed its Claude AI models with high-quality training data. The operation was deliberately kept under wraps, with internal instructions stating, “We don’t want it to be known that we are working on this.” (prensa.com)
These revelations emerged through more than 4,000 pages of court filings unsealed in a copyright lawsuit brought by authors against Anthropic. The documents confirm that the company moved away from using pirated digital libraries like LibGen and Pirate Library Mirror—previously used by co-founder Ben Mann—to build its own physical repository. (prensa.com)
According to the filings, Anthropic’s spending on book acquisitions reached into the tens of millions within approximately one year. The scale of the operation—buying, scanning, and destroying millions of books—underscores the lengths to which AI developers are going to secure high-quality, proprietary training data. (prensa.com)
Anthropic has defended its actions by stating that the scanned books were purchased legitimately and that the resulting digital corpus was used internally—not for commercial resale. The company also emphasized that it ceased reliance on unauthorized digital libraries. Nonetheless, the destructive nature of the process and the secrecy surrounding it have sparked significant controversy. (prensa.com)
The unsealed documents and subsequent reporting raise broader questions about the ethics of data acquisition in AI development, especially when it involves irreversible destruction of physical media. As AI models continue to scale, the tension between data quality, legality, and cultural preservation is likely to intensify.
