Sony and Warner say Anthropic knowingly used pirate libraries to build Claude

Anthropic likes to portray itself as one of the responsible players in artificial intelligence, a company supposedly more concerned with safety and ethics than some of its rivals. A new lawsuit from Sony Music Publishing and Warner Chappell Music paints a much uglier picture, accusing Anthropic and two of its founders of knowingly helping themselves to enormous amounts of copyrighted material while building the technology behind Claude.

The music publishers filed the lawsuit Friday in the U.S. District Court for the Northern District of California against Anthropic, CEO Dario Amodei, and co-founder Benjamin Mann. They accuse the defendants of torrenting, scraping, downloading, scanning, and copying copyrighted material on a massive scale, including musical compositions owned by Sony and Warner.

At the center of the complaint is an accusation that should make anyone who has ever been lectured about piracy raise an eyebrow. According to Sony and Warner, Anthropic turned to pirate libraries because acquiring huge amounts of copyrighted material legally would have been slower, more difficult, and more expensive.

The lawsuit says Mann personally used BitTorrent in June 2021 to obtain at least five million books from Library Genesis, better known as LibGen. Sony and Warner allege Amodei knew about and approved the activity, while Anthropic employees later allegedly torrented another two million books from Pirate Library Mirror, or PiLiMi, in 2022.

That works out to more than seven million allegedly pirated books. The publishers say those collections included songbooks and sheet music containing their copyrighted compositions, including songs such as “Livin’ on a Prayer,” “September,” “Great Balls of Fire,” “Ramblin’ Man,” and “Hallelujah.”

The alleged internal conversations make the situation look even worse for Anthropic. The complaint cites previously disclosed material in which Mann allegedly described LibGen as “sketchy AF,” while Anthropic’s own Archive Team reportedly called LibGen a “blatant violation of copyright.” Sony and Warner say Amodei nevertheless approved Mann obtaining the LibGen material because torrenting it was cheaper than buying the books.

In other words, if the publishers’ account is accurate, this wasn’t a case of Anthropic accidentally stumbling across questionable training data supplied by some mysterious third party. Sony and Warner are alleging that people at the company knew exactly what kind of source they were dealing with and proceeded anyway.

The PiLiMi allegations aren’t much prettier. When Mann learned the pirate library was available for torrenting, the complaint says he sent the link to Anthropic colleagues with the message “[J]ust in time!” Another employee allegedly responded, “zlibrary my beloved.” Engineers then compared the collection with the roughly five million LibGen books Anthropic had already obtained and allegedly torrented another two million books that weren’t duplicates.

Sony and Warner aren’t limiting their accusations to torrenting either. They claim Anthropic scraped copyrighted lyrics from websites that were actually licensed to display them, including MusixMatch and LyricFind, while also obtaining material through datasets such as Books3, The Pile, and Common Crawl.

This is an important part of the publishers’ argument because something being accessible on the web doesn’t necessarily mean anyone can copy it for whatever commercial purpose they want. Sony and Warner specifically argue that licensed lyric websites pay for the right to display their material, and that those arrangements didn’t give Anthropic permission to scoop up the same lyrics for AI development.

Then there is Anthropic’s rather bizarre alleged solution to getting more books. The lawsuit claims the company purchased millions of second-hand physical books, scanned them into digital text, destroyed the books afterward, and placed the resulting text into a massive internal library for potential use in AI training.

Sony and Warner say songbooks and sheet music collections were swept up in this “destructive scanning” operation as well. Perhaps most strikingly, the complaint cites an allegedly internal Anthropic planning document concerning the scanning project that said, “We don’t want it to be known that we are working on this.”

That isn’t the only allegation suggesting Anthropic understood how bad some of this might look. Sony and Warner also accuse the company of intentionally stripping copyright information from material while preparing it for AI training.

According to the complaint, Anthropic evaluated extraction tools that could separate the useful text it wanted from things such as website footers and copyright notices. The publishers point to internal discussions in which one extraction method was allegedly criticized for leaving behind too much “useless junk,” while another was preferred because it did a better job removing footers, copyright-owner names, and notices.

Sony and Warner argue that this wasn’t merely innocent cleanup of messy web pages. They accuse Anthropic of deliberately removing identifying information that could help copyright owners discover that their material had ended up inside Anthropic’s training data.

Claude’s actual behavior is another part of the case. The publishers allege Anthropic’s models have generated verbatim or near-verbatim copyrighted lyrics, pointing to material from previous litigation involving songs such as “Uptown Funk,” “Redbone,” “Stay,” “Scars to Your Beautiful,” “California Gurls,” and “We Belong Together.”

The lawsuit even digs into Claude’s fine-tuning process. Sony and Warner cite records in which an Anthropic model was allegedly asked for lyrics to copyrighted songs and reviewers rewarded accurate output over inaccurate lyrics, which the publishers argue encouraged the model to reproduce copyrighted material.

Sony and Warner also attack Anthropic’s ability to generate supposedly original lyrics. They argue Claude learned how to create those lyrics partly from copyrighted compositions Anthropic had no permission to use, potentially allowing AI-generated music to compete against the very songwriters whose work helped train the system.

Of course, a lawsuit contains allegations, not a verdict, and Anthropic will have an opportunity to dispute this account in court. But if Sony and Warner can substantiate what they are alleging, the image presented here isn’t of an AI startup innocently navigating uncertain copyright law. It is of an enormously valuable company that allegedly knew where pirated material came from, understood the copyright concerns surrounding it, and decided obtaining the data was worth the risk.

The publishers are seeking serious money. They want statutory damages of up to $150,000 for each work found to have been willfully infringed, along with as much as $25,000 for each violation involving the removal or alteration of copyright management information.

Money may not even be the most threatening part of the case for Anthropic. Sony and Warner want a court order requiring the company to account for its training data and methods, identify their copyrighted works used to train its AI models, and explain how those works were collected, copied, processed, and encoded.

They also want Anthropic ordered to destroy infringing copies of their copyrighted works under court supervision and submit a sworn report explaining how it complied. Depending on how broadly a court ultimately interprets such an order, that demand could make this case about much more than writing a very large check.

Anthropic has built an extraordinarily valuable business around Claude while cultivating a reputation for taking AI safety seriously. Sony and Warner are now asking a court to examine what may have happened behind that respectable exterior, and their allegations raise an uncomfortable question for the entire AI industry: how responsible can an AI company really claim to be if it knowingly built its technology with material obtained from pirate libraries?

Anthropic hasn’t lost this case, and Sony and Warner still have to prove their claims. But based on the allegations in this complaint, the publishers aren’t portraying Anthropic as a company that simply misunderstood the complicated rules surrounding AI and copyright. They’re portraying it as a company that knew exactly what it was doing.

Support independent tech journalism

NERDS.xyz is independently owned and operated. If you enjoy my coverage of Linux, AI, hardware, cybersecurity, and tech culture, consider supporting the site on Ko-fi.

Support NERDS.xyz
Written by

Brian Fagioli

Technology journalist and founder of NERDS.xyz

Brian Fagioli is a technology journalist and founder of NERDS.xyz. A former BetaNews writer, he has spent over a decade covering Linux, hardware, software, cybersecurity, and AI with a no nonsense approach for real nerds.

Leave a Comment