According to the complaint, the FBI had shut down LibGen by the end of 2021. However, the publishers allege that online pirates had already copied its content and created a replacement archive known as Z-Library. Although Z-Library was also taken offline, internal messages disclosed during a separate copyright dispute allegedly showed that Anthropic had access to the copied material through the Pirate Library Mirror, or PiLiMi.
The complaint accuses Anthropic executive Jared Mann of directing employees to torrent PiLiMi as soon as it became available. In one message to a co-worker, Mann reportedly said that the mirror had arrived “just in the nick of time.” An Anthropic employee allegedly responded: “Dearest, zlibrary.”
Music publishers say the messages suggest that Anthropic tolerated or encouraged the use of pirated material and relied on torrent networks to obtain new AI training data quickly. The lawsuit further claims that a review of the non-confidential catalogs of LibGen and PiLiMi revealed bibliographic information—including book titles, authors, and ISBNs—indicating that Anthropic torrented at least hundreds of books containing sheet music and song lyrics owned by the publishers.
The publishers say they intend to uncover the “full story” of Anthropic’s alleged torrenting activity through the discovery process in the copyright lawsuit.
Music publishers allege Claude was trained on copyrighted songs
The publishers acknowledge that Anthropic denies using books torrented from LibGen or PiLiMi to train its commercial Claude AI models. However, they argue that Anthropic’s position may depend on how the term “training” is defined. The complaint alleges that additional investigation could show that a commercial Claude model was indirectly trained on pirated works. The publishers claim:
“AI development often includes a ‘pre-training’ phase in which one AI model trains on ‘synthetic data’ created by another AI model or receives other behavioral feedback from another AI model. Anthropic trained at least one commercially released Claude model using synthetic data created by a non-commercial AI model trained on text derived from LibGen and/or PiLiMi. Anthropic has employed at least one non-commercial model trained on text derived from LibGen and/or PiLiMi. LibGen and/or PiLiMi provide at least one commercial Claude model with enhanced feedback.”
The complaint also references recently released legal filings from a separate case involving Anthropic’s use of copyrighted books. Those documents allegedly indicate that Anthropic used books downloaded from LibGen in connection with its AI safety and guardrail research. In one filing, Anthropic witnesses testified that the company stopped training large language models on LibGen material but continued using the dataset to test whether long passages generated by its models matched source texts too closely.
Source: arstechnica.com


