深潮TechFlow|7月 22, 2026 04:06
[Accused of 'Distillation' by Chinese Domestic AI Models, Anthropic Ends Up Paying $1.5 Billion for Pirated Data]
According to Deep Tide TechFlow, on July 22, 2024, multiple authors initiated a class-action lawsuit accusing Anthropic of downloading millions of pirated e-books in bulk from 'shadow libraries' like LibGen and incorporating approximately 482,000 copyrighted works into the Claude training dataset. The court subsequently separated 'model training' from 'data acquisition' in its ruling: AI learning from legally obtained books may fall under fair use; however, downloading and storing works from pirated websites constitutes infringement. Anthropic ultimately agreed to a $1.5 billion settlement, compensating an average of $3,000 per work and destroying the related pirated files.
Ironically, Anthropic had previously criticized some Chinese AI teams for extracting Claude outputs via APIs for model distillation, claiming it violated intellectual property rules. While the two actions are not entirely equivalent under the law, they present a stark contrast: when dealing with human-created works, AI companies emphasize the right of models to learn; yet when it comes to their own model outputs, they demand stronger exclusive protections. The core boundary lies here: AI training itself may not be illegal, but the method of acquiring training data must be lawful. (Jin10)
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink