Google’s Gemini AI faces a class-action lawsuit filed in New York federal court that accuses the company of training its models on copyrighted books and articles without permission. The complaint, brought by a coalition that includes major publishers and author Scott Turow, says Google deliberately stripped copyright information to hide the use of “stolen materials” in its generative-AI system.
The lawsuit and its core claims
The filing in the U.S. District Court for the Southern District of New York lists Hachette, Cencage, Elsevier, the S.C.R.I.B.E. collective and Turow as plaintiffs. They allege Google went beyond the “snippet-only” arrangement that governed Google Books, using full-text works from that archive and from titles uploaded to Google Play to train Gemini. According to the complaint, Google not only scraped the content but also altered or removed copyright metadata, a step the plaintiffs say was intended to conceal the infringement.
An internal Google memo referenced in the suit warns that employing copyrighted books for AI training could be “highly problematic” and expose the company to fines ranging from $10 billion to $100 billion. The plaintiffs point to a recent $1.5 billion penalty against another AI firm for unauthorized data use as a benchmark for the scale of potential liability.
How we got here
Google’s relationship with book publishers began with Google Books, a searchable index that displayed short excerpts and bibliographic details while leaving the full text behind a paywall. That model relied on licensing agreements that limited Google to displaying snippets. Over the years, Google expanded its holdings through the Google Play store, where publishers and authors upload full-text e-books for sale.
The plaintiffs argue that Google’s shift from a “scope-limited” indexing service to a data source for training large language models (LLMs) breaches the original agreements. They say the company never secured the broad permissions needed to ingest entire works into Gemini’s training pipeline.
What’s at stake for Google and the AI industry
If a court finds that using copyrighted material without a license violates copyright law, the decision could reshape how AI developers assemble training data.
Possible defenses and counterpoints
Google will likely lean on the “fair use” doctrine, which permits limited use of copyrighted material for commentary, criticism, or transformation. Recent California rulings have treated AI training as a transformative activity, granting it fair-use protection. The New York plaintiffs filed outside that jurisdiction to avoid those precedents.
The plaintiffs counter that removing copyright metadata shows an intent to conceal the source, undermining any claim of good-faith transformation. They also cite the internal memo as proof that Google recognized the legal risk.
What to watch next
The case will move through pre-trial motions that could shape the arguments available to both sides, including whether the court will apply the fair-use analysis used in California or adopt a different standard. A ruling on jurisdiction could determine whether New York courts become the primary venue for AI-related copyright disputes.
Bottom line: The New York suit against Google’s Gemini could redraw the line between permissible data use and copyright infringement, forcing AI developers to rethink how they acquire the raw material that powers their models and reshaping the economics of the entire industry.
