Restricting fair use for AI training would raise costs, suppress competition, and widen the knowledge divide.
For more than two years, opposing armies have been massing and digging in as they prepare for a battle royale to decide the future of their respective domains.
No, the Iron Throne doesn’t hang in the balance, but the future of generative AI, and its potential to better the lives of literally everyone everywhere, does.
Those are the stakes in a massive consolidated copyright legal case against OpenAI brought by dozens of authors, including George R.R. Martin and major news organizations. After years of preliminary jousting by lawyers and experts, the central legal issue will come to a head in the fall and is expected to be argued early next year before a federal judge in New York City.
The key issue is “fair use,” a cornerstone of American copyright law since before the Civil War and written into the Copyright Act by Congress 50 years ago this fall. It expressly allows copyrighted works to be used without the prior permission of their owners to create something fundamentally new and different, provided the new work does not substitute for the original itself.
Ironically, authors and news outlets regularly rely on “fair use” of copyrighted material to write their own books and articles. Yet that hasn’t stopped many of them from going to court to stop generative AI developers from using copyrighted material, much of it from the open internet, to initially train large language models like ChatGPT.
Plaintiffs are asking in many cases for court orders that would force developers to pull their existing gen AI models off the market entirely until they can be retrained using only licensed material.
Developers like OpenAI push back with a different reading of the law. They argue that their use of copyrighted material is a textbook case of fair use under decades of precedent, because the material on which gen AI models are trained isn’t copied, stored or resold – it’s transformed into a new kind of tool, not a substitute for any specific book or article. They also cite a clear line of Supreme Court and appellate rulings that the maker of a multi-purpose product isn’t usually on the hook when someone else uses that product to violate copyright law.
For most of the general public, this may sound like an obscure legal fight between powerful companies over money, but it’s much bigger than that. If courts dramatically narrow fair use for AI training, the consequences would reach every classroom, laboratory, startup, library, marketplace and workplace that increasingly relies on generative artificial intelligence. That could negatively affect every one of us in at least four important ways:
1. Innovation would become harder, and the next generation of AI companies might never get started.
Today’s leading AI models require enormous amounts of information during training. Instead of using fair use material, the alternative would be to buy or lease that material. It sounds simple enough, but the volume of material required is astronomically large and would be astronomically expensive. That’s a non-starter for start-ups. Worse yet, the cost and mechanics of simply identifying and negotiating for the huge amount of data necessary would itself be prohibitive for all but the biggest existing developers.
2. Generative AI would become more expensive and less accessible.
Requiring gen AI developers to use only licensed training material would invariably increase the cost of developing new models. These costs would ultimately flow to users. That would affect not only commercial enterprises, but schools, libraries, other public institutions and governments of every size, as well as individual consumers. As developers unable to absorb all of the substantially heightened training costs struggle to compete, all users should expect higher subscription costs, and likely lower usage limits.
The largest corporations would still be able to afford premium AI systems. Smaller businesses, nonprofits, and public institutions would have fewer choices and fewer capabilities.
3. Scientific research, technical innovation, and knowledge discovery would slow.
If access to training data became substantially more restricted, specialized research models often built from general-purpose gen AI models would likely become more difficult and more expensive to build. The pace of scientific progress would slow, not because researchers lacked ideas, but because they lacked affordable tools. The public’s access to knowledge would be limited and more fragmented.
Perhaps the least appreciated consequence of eliminating fair use for AI training is how it would limit and fragment the knowledge available to everyone through future AI systems. Today’s best models owe their benefits directly to training on extraordinarily broad collections of books, newspapers, academic writing, historical documents, and countless other sources. If every copyright owner could decide whether its works were used in gen AI training, future models would inevitably become patchworks with a fraction of their former power and utility.
In that world, some models might include particular newspapers but not others, contemporary books but not historical archives, and major commercial publications but not local and independent reporting.
For ordinary users, that could mean:
- Less complete answers about current events, history, literature and specialized subjects;
- Greater dependence on paywalled material and multiple subscriptions;
- Potentially varied answers to the same inquiry depending on the user’s employer, school or subscription level;
- Reduced coverage of local, niche, minority-language and out-of-print material; and
- A widening gap between well-funded and underfunded educational institutions.
Here there be dragons
Ancient maps of the globe advised navigators to avoid uncharted waters by marking them with the often appropriately illustrated warning: “Here there be dragons.” The plaintiffs in these cases ask courts to believe that gen AI has taken us into exactly that kind of unexplored and dangerous legal territory.
The fact is, we have a clear map that has provided copyright holders with strong rights and given all of us maximal benefit from incredible new technologies like generative AI. That map is the Constitution, which says unambiguously that the purpose of copyright law is “to promote the progress of science and useful arts.”
Robust fair use has been and remains the principal way U.S. copyright law strikes the balance between compensation and innovation. That must continue. The real danger to copyright isn’t from generative AI. It’s posed by fire-breathing plaintiffs intent on impeding a new technology’s development by incinerating the foundations of fair use.
/>i
