Lawsuit says NVIDIA looked to Anna’s Archive for AI training books
A new court filing claims NVIDIA tried to secure access to a large stash of pirated books to train its AI models.

The amended class-action complaint filed in the U.S. District Court for the Northern District of California claims NVIDIA staff reached out to Anna’s Archive, a site that hosts millions of pirated books and academic papers.
The filing says the talks were about getting “high-speed access” to the archive’s data. Anna’s Archive allegedly told NVIDIA the material was obtained illegally and asked if NVIDIA had internal approval to keep going. The complaint claims management signed off soon after.

Source: Torrent Freak
Anna’s Archive is said to have offered access to about 500 terabytes of data. That collection allegedly included millions of books, some of which are normally available only through the Internet Archive and its controlled digital lending system. The filing does not say whether NVIDIA paid for the access or used the data that was offered.

Source: Torrent Freak
The authors accuse NVIDIA of using other pirate sources, including the Books3 dataset and sites such as Library Genesis, Sci-Hub, and Z-Library. Another claim says NVIDIA provided scripts or tools that allowed customers to download parts of “The Pile” dataset, which includes Books3 (a large dataset that contains 200K books). The authors argue this led to contributory and vicarious copyright infringement, since customers could access pirated books through NVIDIA-provided tools.

Source: Torrent Freak
NVIDIA has argued before that AI training is fair use and that models learn patterns instead of storing books. The case is ongoing, and these details come from the plaintiffs’ latest filing.
Source: Torrent Freak