Policy & Regulation
Google faces new AI-training copyright lawsuit from publishers
A group of publishers and authors, including Hachette and Scott Turow, has filed a class action lawsuit accusing Google of training its Gemini AI models on their copyrighted works without permission.
A group of publishers and authors has filed a class action lawsuit against Google, accusing the company of using their copyrighted works to train its AI platform, Gemini. The plaintiffs — who include Hachette, Cengage, Elsevier, author Scott Turow, and S.C.R.I.B.E. — also allege that Google intentionally removed or changed copyright information on these works to conceal that, in the lawsuit’s characterization, its Gemini models had been trained on stolen material.
The suit is one of many complaints publishers, authors, and other copyright holders have filed against AI companies including Google, Meta, OpenAI, and Anthropic. Many of these cases are still pending, though two early court decisions in California have favored the AI companies, ruling that using copyrighted works for AI training counts as “fair use” under U.S. copyright law. Those California rulings don’t necessarily bode well for how other courts may view AI companies’ fair use defense, though the legal questions are considered too nuanced for the decisions to set a clear precedent. Anthropic, however, was fined $1.5 billion for pirating the works it trained on — the largest payout in the history of U.S. copyright law — with around half a million writers eligible for payments of at least $3,000, though many opted out to pursue further legal action.
The Google suit was filed in the U.S. District Court for the Southern District of New York, giving a different judge a chance to weigh in. Unlike some other cases, the publishers here describe a longer, more nuanced relationship with Google: the lawsuit says publishers and authors have long supplied copyrighted works so Google Books could make them searchable, showing only short snippets and bibliographic information rather than full text. The plaintiffs claim Google trained Gemini on copies of these books, as well as on books uploaded to the Google Play store, without ever receiving permission to do so.
The plaintiffs also cite an internal Google document that allegedly said using copyrighted books for AI training could be seriously problematic for the company and might result in “$10Bs-$100Bs in potential fines.” Google did not immediately respond to a request for comment.
Why it matters
It adds Google to a growing list of AI companies fighting copyright claims over training data, at a moment when courts remain divided on whether “fair use” shields large-scale AI training — and Anthropic’s $1.5 billion settlement shows how costly losing that argument can be.