
Song lyrics have become the latest flashpoint in the legal struggle over artificial intelligence and copyright. Sony Music Publishing and Warner Chappell have sued Anthropic in a Northern California court over lyrics allegedly copied from pirate archives and used to train Claude, the company’s family of large language models. The lawsuit is notable not only for the size of the damages demanded but for the fact that it names Dario Amodei, Anthropic’s chief executive, and Benjamin Mann as individual defendants. In the complaint, the publishers describe a “brazen campaign of illegally torrenting, scraping, and downloading copyrighted works on a massive scale.”
What the complaint alleges
The publishers claim that Anthropic collected millions of songs from online sources without obtaining licences from rightsholders. According to the complaint, those sources included Library Genesis and Pirate Library Mirror, two shadow libraries that are widely used in academic circles but host material without authorisation. The same archives were central to an earlier set of legal claims against Anthropic. That fight ended in a $1.5 billion settlement with authors who accused the company of using their books without permission. The music publishers now want their own relief, but the per-work figure they are seeking is far above what authors received in that deal.
Under the US Copyright Act, a plaintiff can choose statutory damages instead of proving actual harm. For wilful infringement, the ceiling is $150,000 per work. The publishers have asked for that maximum amount for each composition allegedly used in training. A jury trial has also been requested. The complaint lists several well-known songs, including Survivor’s Eye of the Tiger, Leonard Cohen’s Hallelujah, Earth, Wind and Fire’s September, Bon Jovi’s Livin’ On a Prayer and Jerry Lee Lewis’s Great Balls of Fire. Compositions by Mariah Carey and Taylor Swift are also said to be among the tracks found in the training data.
The choice of such famous songs appears deliberate. It makes the harm easy for a jury to understand. These are not obscure line items in a database; they are lyrics that millions of people know by heart. If Claude can reproduce lines from those songs, the publishers argue, it is because the model has memorised protected expression rather than learned a general pattern about language.
Why individual defendants were named
Naming Amodei and Mann personally is a significant legal tactic. Corporate veil protections generally shield employees from liability for actions taken on behalf of a company, but copyright law allows direct liability for anyone who participates in infringement. By placing the executives in the suit, the publishers may be trying to force disclosure of internal decisions about training data. They may also be trying to raise the stakes for Anthropic’s leadership. If the case goes to trial, the founders could be asked about what they knew, when they knew it, and whether anyone in the company understood the potential scale of the infringement.
Anthropic has not yet filed a substantive answer. Its likely defences include fair use and the argument that training on copyrighted material is transformative. The company may also argue that song lyrics are not meaningfully reproduced in ordinary conversations with Claude, and that the specific outputs at issue were produced through prompting designed to draw out trademarked or copyrighted text.
The gap between demand and settlement
One of the most striking numbers in the dispute is not the $150,000 ceiling. It is the contrast with the previous settlement involving books. In that case, authors received about $3,000 per title, which was then split with their publisher. That left roughly $1,500 on each side. The difference between those figures and the publishers’ opening demand shows how far apart the two sides are. One is an agreed compromise; the other is an initial ask in an unanswered lawsuit. The final number, if any settlement is reached, could land somewhere between the two, but only after discovery reveals how many songs were actually used and in what manner.
For the music industry, the stakes are larger than one lawsuit. Songwriters rely on licensing revenue for their livelihoods. If an AI company can train on lyrics without paying, then a foundation of the music business is called into question. That is why the publishers are pressing for statutory damages rather than a reasonable licence fee.
A German court has already weighed in
Across the Atlantic, a court in Munich has already addressed the same question in a case against OpenAI. In November 2025, the Regional Court of Munich ruled that memorising lyrics inside a model is a reproduction under copyright law. It also found that outputs reciting those lyrics amount to communication to the public. The ruling said the EU’s text and data mining exception did not apply, because permanent memorisation goes beyond transient analysis and because the rightsholder had opted out. The judgment is not final, but it offers a powerful precedent for publishers litigating elsewhere.
The European exception for text and data mining has two important limits. The first is that it only applies to works the miner had lawful access to. A pirate library is never lawful access. The second is that rightsholders may opt out, so even materials found on legitimate sites can be off-limits if the owner has reserved rights. The Munich ruling suggests that, at least in Europe, training data sourced from unauthorised archives is not protected by a technical exception. That is directly relevant to Anthropic because the complaint identifies Library Genesis and Pirate Library Mirror as the alleged sources.
Transparency under the AI Act
The EU has gone further with the AI Act. General purpose model providers are required to maintain a copyright policy and publish a summary of the content used in training. The law is enforced by a dedicated unit in Brussels. This creates a striking asymmetry between American and European rightsholders. In the United States, a publisher may need to file a lawsuit and go through discovery to learn what was actually used. In Europe, the law gives rightsholders a direct entitlement to information. The same company might therefore have very different obligations depending on where a claim is brought.
This asymmetry may push more disputes into Europe. The Munich ruling shows that European courts are willing to take a strict view of model memorisation, and the AI Act provides a transparency mechanism that does not exist in the US. Music publishers have global catalogs and can choose where to litigate. A favourable ruling in one jurisdiction can influence negotiations everywhere, especially when the same training data is used in models distributed internationally.
The Anthropic lawsuit is still in its early stages. The publishers have made a high-stakes opening move by naming executives and demanding the maximum statutory damages. The company has yet to answer, and the discovery process could take years. But the legal environment is shifting quickly. Between the German court’s decision, the AI Act, and this new case, the question of whether lyrics may be used as training data is moving from abstract policy debate to concrete courtroom conflict.
Source:TNW | Anthropic News
