Lawsuits alleging the makers of AI models infringed on copyrights during the training process are piling up. A new proposed class action lawsuit brought against Meta by major publishers in a Manhattan federal court will see a first hearing in September. Previously, groups of authors had sued tech companies, but in the cases of another suit against Meta and one against Anthropic had not succeeded in showing that their copyrights had been breached by the training process itself.
However, the latter case found that Anthropic had stored a whopping 7 million pirated books to train Claude, which was ruled illegal. The finding also shows the sheer volume of works that were fed to large language models.
An analysis by data scientist Olivier Khatib, CEO of T1U.ai, shows that between five major LLMs, Meta's Llama is estimated to have been trained on approximately a quarter of book material as of 2025, actually much higher that the estimated share for Claude. Comparably many books are also thought to have been used on the training of ChatGPT and Google's Gemini. The calculation based on company disclosures, model documentation and additional analysis shows that Claude used a high share of academic papers, while Gemini was trained on a majority of web crawl data. Grok by Elon Musk's company xAI, maybe predictably so, used a higher share of social media input.





















