Monday, August 3, 2026

Policy & Regulation

Encyclopedia Britannica sues OpenAI over copyright infringement

Encyclopedia Britannica and Merriam-Webster have sued OpenAI, alleging "massive copyright infringement" regarding the use of their content in LLM training and RAG workflows.

Encyclopedia Britannica sues OpenAI over copyright infringement

Encyclopedia Britannica and its dictionary subsidiary Merriam-Webster have filed a lawsuit against OpenAI. According to the complaint, the publishers allege that the artificial intelligence company has committed “massive copyright infringement.” The publisher asserts that it retains the copyright to nearly 100,000 online articles, which it claims OpenAI scraped and used to train its large language models without obtaining permission.

In the legal filing, the publishers target both the training of OpenAI’s models and its live search features. Britannica alleges that OpenAI violates the Lanham Act—a U.S. federal trademark statute—when its models generate false information, known as hallucinations, and attribute them falsely to the publisher. Furthermore, the lawsuit alleges that OpenAI’s use of these articles in ChatGPT’s Retrieval Augmented Generation (RAG) workflow—a technique used to enhance responses with external data—directly harms their business. The legal filing states that “ChatGPT starves web publishers like [Britannica] of revenue by generating responses to users’ queries that substitute, and directly compete with, the content from publishers like [Britannica],” while also warning that hallucinations jeopardize public access to trustworthy information.

The legal action adds to a growing list of copyright disputes between AI developers and content creators. OpenAI is already facing similar lawsuits from other major publishers, including The New York Times and Ziff Davis, alongside more than a dozen newspapers across the U.S. and Canada. Currently, legal precedent has not clearly established whether using copyrighted content to train a large language model constitutes copyright infringement. However, a separate case involving AI lab Anthropic serves as a key comparative benchmark for this legal uncertainty. In that dispute, federal judge William Alsup ruled that Anthropic was “illegally downloading millions of books” rather than paying for them, which ultimately led to a $1.5 billion class action settlement for the impacted writers.

Why it matters

This lawsuit represents a significant escalation in the ongoing conflict between legacy publishers and AI labs, specifically targeting how RAG tools and training data impact the business models of content creators.