Artificial Intelligence
Llama AI Develops Massive Hunger for Data Points
Major publishing houses have sued Meta in a proposed class action lawsuit for alleged copyright infringements during the training of its AI model Llama. The group of publishers and an author sued the company in early May saying Meta used their proprietary books and journal articles for AI training without permission. Many of the companies – Elsevier, Cengage, Hachette as well as Macmillan and McGraw Hill – focus on educational and academic publishing. A first hearing is set for September 15 in a Manhattan federal court.
Data from Epoch AI shows how data-hungry the latest AI models have become. In the case of Meta's Llama, a first version released in February of 2023 used 1.4 trillion unique data points. By mid-2024, this had rapidly increased to 15 trillion before finally reaching a whopping 30 trillion as of the latest versions in April 2025. Model DeepSeek out of China went a similar route, starting in June 2024 at 3.2 trillion data points and having increased that to 14.8 trillion by the release of the next version in early 2025.
The case is just the latest in a string of lawsuits against tech companies training AI models. Dozens of authors, news outlets, visual artists and others have so far sued with similar complaints, according to Reuters. One of the more widely reported suits is that of comedian Sarah Silverman and authors Richard Kadrey and Christopher Golden filed in San Francisco, also against Meta. The group suffered a first defeat in June, however, as a U.S. district judge found that they had not successfully laid out how Meta had processed their published works beyond fair use clauses and into infringement territory. The authors appealed. A suit against Anthropic, makers of AI model Claude, similarly found that the training of large-language models using copyrighted works was transformative enough to fall under the fair use doctrine. The case, however, also determined that the company simply storing pirated versions of more than 7 million books in a database was illegal and would have resulted in a fine (which was in the end paid out via a settlement).
Description
This chart shows the Number of unique data points used in training datasets for Meta's Llama AI models.
Related Infographics
Any more questions?
Get in touch with us quickly and easily.
We are happy to help!
Statista Content & Design
Need infographics, animated videos, presentations, data research or social media charts?