NEW YORK / RankWire.AI / – Three leading U.S. publishers have filed a lawsuit against Google, accusing the tech giant of copyright infringement related to its Gemini artificial intelligence platform. Hachette Book Group, Cengage Learning, and Elsevier initiated a proposed class action alongside author Scott Turow and his firm, S.C.R.I.B.E. The complaint was lodged on July 10 in a federal court in New York, claiming Google unlawfully copied millions of copyrighted books and journal articles during the development and training of Gemini models without permission.

According to the plaintiffs, Google acquired the copyrighted material through Google Books, Google Play Books, and Google Scholar. Publishers and authors had submitted their works to these platforms to support search functionalities, sales, and research activities, the complaint states. However, the filing argues that these arrangements did not grant Google the right to reproduce the works for commercial AI training purposes. Additionally, Google is accused of utilizing web-scraped datasets that incorporated content from pirate sites and subscription-based services behind paywalls.
Google is facing four allegations outlined in the 57-page complaint. Three of these claims involve illegal reproduction via Google services, web scraping, and the process of developing or training Gemini. The fourth alleges violations of the Digital Millennium Copyright Act, asserting that Google removed or altered copyright management information, including author names, ownership details, and publication data. As of July 15, the court had yet to rule on these claims or grant class-action certification.
Four allegations focus on Gemini’s training data
The proposed class encompasses owners of registered U.S. copyrights in books and journal articles. To qualify, works must have an International Standard Book Number (ISBN) or a Digital Object Identifier (DOI) or an International Standard Serial Number (ISSN). This definition includes works Google allegedly copied from its platforms, downloaded via web scraping, or reproduced during the creation of Gemini. Eligibility is also limited to works registered within the deadlines specified in the complaint.
The complaint cites examples from Hachette, Cengage, and Elsevier, including fiction, textbooks, and scholarly publications, as instances of the alleged unauthorized copying. It also references internal Google evaluations concerning legal risks associated with publisher-provided content. One such assessment reportedly warned of potential fines between $10 billion and $100 billion, according to the plaintiffs. The court has not issued any findings regarding these internal documents.
Claims for damages and transparency sought by plaintiffs
The plaintiffs are seeking statutory or actual damages, along with profits derived from any proven infringement. They also request an injunction, coverage of legal expenses, and a jury trial. Their proposed order would mandate Google to disclose the sources and collection methods used to train Gemini. Additionally, they ask the court to oversee the destruction of any unauthorized copies under Google’s control. The complaint does not specify a total dollar amount for the damages sought.
This lawsuit in New York follows an earlier effort by Hachette and Cengage to join separate Google AI copyright litigation in California. The Association of American Publishers noted that the new case preserves claims outside the scope of that proceeding. The current suit involves Elsevier, Turow, and S.C.R.I.B.E. alongside the two publishers, and seeks to determine whether Google’s Gemini training procedures and data collection practices violate federal copyright law and the Digital Millennium Copyright Act.
