The U.S. Department of Justice has issued a significant legal intervention in the ongoing copyright war between media giants and AI developers. By arguing that training Large Language Models (LLMs) on copyrighted data constitutes "fair use," the DOJ has provided a massive legal shield for companies like OpenAI and Microsoft.
The DOJ’s Argument: Training vs. Output
At the heart of the consolidated lawsuit involving The New York Times is a fundamental distinction between how an AI learns and what it produces. The New York Times alleges that millions of its articles were used without permission to train models such as GPT-4, causing billions of dollars in damages and creating products that directly compete with the newspaper.
However, the DOJ’s recent filing argues that copyright infringement occurs at the point of output, not the point of ingestion. The department contends that while entire works are copied during the training phase, these copies are never made publicly available. Furthermore, the DOJ asserts that the outputs of models like GPT-4 "often if not always lack substantial similarity" to the original source material. By separating the training process from the generated content, the DOJ is attempting to decouple the act of machine learning from the legal concept of market substitution.
The "Hemingway Analogy" and Human Creativity
To make the concept of machine learning relatable, the DOJ invoked a literary analogy involving author Joan Didion. The filing noted that as a teenager, Didion would copy Ernest Hemingway’s stories to analyze his sentence structure and learn his craft. The DOJ argues that if the law treated machine training as infringement, a writer like Didion could theoretically face liability every time she published new work, as her learning process would be inextricably linked to her subsequent writing.
The department further argues that imposing strict liability for AI training would paradoxically stifle the very creativity copyright law is intended to protect. With human beings increasingly using LLMs to draft and edit original works, the DOJ suggests that making training impermissible without massive licensing schemes would cripple AI-driven innovation.
Challenging the US Copyright Office
The DOJ’s stance directly contradicts recent findings from the US Copyright Office. Former Register Shira Perlmutter had previously argued against a blanket "fair use" defense, noting that AI operates at a scale and speed far beyond human capability and often creates commercial products that compete directly with the original content creators.
The DOJ has gone on the offensive against this assessment, stating that Perlmutter’s report carries no binding legal authority and fails to account for established case law regarding case-by-case analysis. The department warns that imposing broad liability would effectively mandate a licensing regime that would render the development of advanced AI models legally and financially impossible.
Key Takeaways
- Separation of Training and Output: The DOJ argues that copying text for training does not constitute infringement if the resulting model outputs do not show "substantial similarity" to the original works.
- Protection of Innovation: The department warns that requiring licenses for all training data would stifle creativity and prevent the development of LLMs that assist in human content creation.
- A Legal Bellwether: This intervention creates a high-stakes conflict between the DOJ and the US Copyright Office, setting the stage for a definitive court ruling on the future of generative AI.
ARTICLE: The U.S. Department of Justice filed an amicus brief in the New York Times’ copyright lawsuit, arguing that copying protected text to train large language models (LLMs) qualifies as fair use. The filing gives companies such as OpenAI and Microsoft a legal shield that could keep them out of a damages pool that the newspaper estimates in the billions.
Why the case matters now
The lawsuit consolidates several claims that AI developers fed millions of newspaper articles into their models without permission, then released products that compete directly with the Times’ own reporting. If a court rejects the DOJ’s fair-use argument, AI firms could face massive licensing bills or be forced to halt development of the next generation of conversational agents.
The DOJ’s core argument
У документі робочий процес ШІ розділено на два окремі етапи. По-перше, модель поглинає великі масиви текстів; по-друге, вона генерує відповіді на запити користувачів. Департамент стверджує, що порушення виникає лише тоді, коли об'єкт авторського права відтворюється у спосіб, доступний для громадськості. Оскільки копії для навчання ніколи не залишають серверів розробника, сам акт поглинання не досягає цього порогу.
DOJ також спирається на тест «суттєвої схожості» — давній стандарт авторського права. Міністерство стверджує, що текст, який генерують такі моделі, як GPT-4, «часто, якщо не завжди,» настільки відрізняється від будь-якого окремого джерела, що позивач не може довести необхідну схожість. Коротше кажучи, у документі аргументується, що юридичний акцент має бути на кінцевому результаті, а не на внутрішньому процесі навчання.
Літературна аналогія для ілюстрації тези
Щоб зробити технічний аргумент зрозумілішим, у поданні наводиться історія про підлітка-письменника, який копіював оповідання Ернеста Гемінґвея, щоб вивчити його стиль. DOJ зазначає, що якби закон розглядав таку вправу з навчання як порушення, на письменницю могли б подати в суд щоразу, коли вона публікувала б новий твір, навіть якщо він був оригінальним. Департамент попереджає, що поширення такої ж логіки на ШІ «паралізує саме те творче начало, яке закон про авторське право мав захищати».
Ставки для розробників ШІ та творців контенту
- Ризик для інновацій: Вимога ліцензувати кожен фрагмент тексту, використаний для навчання, може призвести до настільки високих витрат, що створення передових моделей стане фінансово нерентабельним.
- Творчий робочий процес: Люди-письменники дедалі частіше покладаються на LLM для написання чернеток, редагування та мозкового штурму. Якщо процес навчання буде визнано незаконним, ці інструменти можуть зникнути, що змінить способи створення контенту в різних галузях.
- Вплив на ринок: Газетна індустрія стверджує, що моделі ШІ, які можуть відповідати на запитання або генерувати тексти, схожі на новини, підривають цінність оригінальної журналістики, потенційно відтягуючи рекламні доходи та підписки.
Контраргумент Бюро авторського права
Нещодавній звіт Бюро авторського права США, підготовлений за часів колишньої Реєстраторки Шири Перлмуттер, заперечує можливість використання загального захисту на основі «добросовісного використання» (fair use). Бюро виділило три фактори, які відрізняють ШІ від людського навчання: величезний обсяг скопійованого матеріалу, швидкість, з якою він обробляється, і той факт, що отримані моделі можуть бути упаковані та продані як комерційні продукти, які безпосередньо конкурують із творцями першоджерел.
DOJ спростовує ці тези, зазначаючи, що звіт не має обов'язкової юридичної сили і що чинна судова практика вже вимагає аналізу кожного факту на предмет добросовісного використання. Департамент попереджає, що встановлення «широкої відповідальності» фактично змусить галузь перейти до режиму ліцензування, який «зробить розробку передових моделей ШІ юридично та фінансово неможливою».
Що може статися далі
Окружний суд, який розглядає справу Times, має зважити позицію DOJ щодо добросовісного використання та висновки Бюро авторського права. Рішення на користь DOJ створить прецедент, згідно з яким дані для навчання можна збирати без прямого дозволу, за умови, що результати роботи моделі не мають суттєвої схожості з оригіналами. Рішення на користь газети може спровокувати хвилю переговорів про ліцензування, що потенційно змінить економіку досліджень у сфері ШІ.
Підсумок
Подання DOJ перетворює питання про те, чи є навчання ШІ «добросовісним використанням», на битву в суді з високими ставками. Результат визначить, чи зможуть розробники продовжувати створювати потужні мовні моделі на основі наявних текстів, чи вони муситимуть шукати дорогі ліцензії на кожен фрагмент матеріалу, який вони поглинають. Це рішення відгукнеться в технологічному секторі, медіаіндустрії та ширшій креативній економіці.
