Could a 'Learnright' Be the Solution to the AI vs. Copyright Question?
A new paper makes the case
At least 128 cases involving AI and copyright have been filed in the U.S., according to the tally kept by Prof. Edward Lee on his invaluable ChatGPTIsEatingTheWorld blog, of which 116 are currently active. The remaining dozen were dismissed, settled or remain on appeal. Another 162 cases have been filed outside the U.S.
The largest number of U.S. cases—50 of the 128—have been filed by authors or book publishers, followed by news publishers (18), music publishers or record labels (17), video and film producers (15), and visual artists (12). OpenAI has been on the receiving end of the largest number of those cases, at 22, followed by Meta (16), Microsoft (12), and Anthropic (11).
The vast majority of those cases concern the alleged unlicensed use of copyright-protected works to train generative AI models, which the plaintiffs argue is infringing on their exclusive rights. In the vast majority of those, the AI company defendants have raised a defense of fair use, maintaining that the works included in their training datasets were made publicly available to read by the rights owner and that the process of training an AI model is transformative and does not implicate any of the five defined rights included in the copyright bundle, and therefore is fair use.
At the core of all of those legal disputes is the question of who should get paid for and from the fruits of generative AI. Leaving aside for now Ed Zitron’s argument that no one currently is reaping any fruit from AI, and may never, under the current structure of the industry, AI developers are benefitting (if not profiting) from their use of copyrighted works to train their models while creators and rights owners, by and large, are not. What the public is gaining or losing is a matter of debate.
Unacknowledged in all those cases, however, is the question of whether resolving the fair use vs. infringement debate would—or even could—actually resolve the question of who gets paid. That’s the view advanced by Profs. Frank Pasquale of Cornell Law School and Thomas W. Malone of MIT Sloan School of Management, and Andrew Ting, a lecturer in law at George Washington University, in a provocative paper published last fall in the Journal of Technology and Intellectual Property.
The paper runs through the pro and con arguments to both sides of the debate and the objections to each, weighing them all against both the law and equity. It concludes that the law, as currently written and applied by courts, cannot fairly balance the equities involved in generative AI training, no matter how the current lawsuits are decided, because it cannot support a viable system of licensing that would allow for a market-based solution.
Instead, the authors propose adding a new independently licensable right to the current copyright bundle that they refer to as the learning right, or “learnright.”
“Thoughtful policy demands respect for the humans whose works are now fueling this surfeit of computerized output,” the authors write. “Given that the myriad copyright cases now slowly making their way through courts may never vindicate these creators, or may do so inadequately, there is a strong normative case for policymakers to support a “learnright” to supplement existing intellectual property protections.”
Since their paper was published, Malone and Pasquale have also published condensed versions of their argument in non-scholarly outlets, including here and here.
As described by Malone, et. al., the learnright would consist of an exclusive right governing machine learning from protected expressive works. It would not extend to any form of human learning.
Unlike the automatic protection provided to the fixation of expression, the authors propose a system of opt-in registration for creators to invoke the new right, somewhat analogous to patent and copyright registries. Registration would identify the rightsholder and provide sufficient information for AI companies to locate the owner and negotiate payment. Any work registered would require permission to be used to train AI models. Developers would also be required to disclose the contents of their datasets.
The authors assume that many, perhaps even most, creators whose work is not likely to have meaningful commercial value as training data would not bother to register their works. That would reduce transaction costs for AI companies by eliminating the need to clear every work they potentially could encounter, allowing them to license a curated set of the most valuable registered data.
Such a statutory structure, the authors contend, would allow for the development of a market-based licensing system based on blanket licensing , without the need of a compulsory license or government-set fees. They envision an ecosystem comprised of four components:
Creators/rightsholders: authors, artists, publishers, news organizations and other content owners.
AI companies: entities training and operating generative AI systems.
Users: consumers and businesses using AI outputs.
Agents/content brokers: intermediaries aggregating and licensing rights.
They envision the fourth stakeholder—agents and content brokers—as analogous to music PROs, offering blanket licenses for a particular category of works and responsible for dispensing the fees collected to rights holders.
“By aggregating various kinds of content from multiple sources, these Agents can greatly simplify the amount of negotiating that needs to go on in this market,” they write. “For example, one Agent might represent many of the major news publishers, another might represent sets of fiction authors (or their publishers), and so forth.”
The concept is not entirely new. Other analysts and commentators have proposed creating some sort of AI training right, either sui generis or as an addition to existing copyright law. But Malone, Pasquale and Ting deserve credit for offering a fully developed proposal for getting past the fair use vs. infringement defile by envisioning a legal structure to support a market-based licensing system that doesn’t require settling the question definitively.
Their proposed system would still face challenges, however. The first is definitional. How exactly would the legally covered act of “learning” be defined such that it doesn’t sweep up currently legal activity? The paper does not include any suggested statutory text.
The harder problem would be antitrust. If there is a single entity offering a license to content from all major news outlets, or all major book publishers, the potential for abuse would be significant. Without competition there would be nothing to restrain the price of a license absent government intervention. ASCAP and BMI were forced to operate under court supervision for decades due to exactly that problem.
There is also no discussion in the paper about how startups or open-source AI developers would fare under the proposed regime. They potentially could be priced out of the market for the most valuable data.
Depending on how the market ultimately shakes out the proposed structure could also face the inverse problem of monopsony, in which a small number of large buyers would be able to dictate prices, preventing rights holders from realizing the full value of their rights.
The paper is also silent on whether retraining, fine-tuning or retrieval-augmented generation would trigger the right. Nor does it specify whether the envisioned licensed would be for one-time use in training, or for a model’s ongoing use of its “learning” as it generates new outputs. If the latter, it’s unclear what would happen when the licensed has expired or been revoked. Would developers be forced to retrain their models without the expired content?
Also not clear is how the authors envision how large, multinational datasets would be handled across jurisdictions, including where the “learnright” has not been adopted.
Some of those challenges, of course, can likely be met through stakeholder negotiation and some clever lawyering. The bigger challenge would be political: actually enacting the necessary amendment to the Copyright Act.
The deeper potential challenge is conceptual. Locating the learnright within the domain of copyright may make a certain intuitive sense. But it’s a bit of odd fit. Unlike the reproduction, distribution, performance, display and derivative work rights, all of which are agnostic as to the method or technology used to exploit them, the envisioned learnright is use-case specific. It applies to the use of otherwise copyrighted works for the purpose of machine learning and in no other circumstance or technological context.
The closest analogy would be the digital audio performance right introduced in 1995. It applies to the public performance of non-dramatic music recordings via digital transmission, but not to the performance of the same recordings via over-the-air broadcast. In that case, the difference in application between use cases reflects the relative political influence of broadcasters and digital platform providers at the time of passage, not the balance of equities or any underlying principle related to the progress of useful arts and sciences. But political influence waxes and wanes with shifts in the market and culture.
The learnright might have a better claim on the balance of equities, but unlike the digital audio performance right it could also hold the potential for conflict with other elements of copyright. Should the Supreme Court ultimately decide that the use of copyrighted works to train AI models qualifies as fair use, for instance, would the learnright override that judgment, or be precluded by it? It’s not clear.
Careful legislative drafting might be able to avoid such potential conflicts to wedge the learnright into the Copyright Act, but it won’t be easy. And pity the first courts asked to draw the line.
To be clear, none of those caveats is intended as criticism of Malone, Pasquale and Ting. As noted, they deserve credit wrestling seriously with the complexity of the problem and offering a serious proposal to resolve it. Their paper is highly recommended reading. What it mostly reveals, however, is just how complex and difficult the challenge generative AI is to copyright.

