Publications

Licensing Training Data and Attributing Copyright of Derivative Content From Large Language Models Can Resolve Up- and Downstream Copyright Issues

Brian ZhouSritan Motati
ICML Workshop on Generative AI and the Law, 2023
BibTeX
Licensing Training Data and Attributing Copyright of Derivative Content From Large Language Models Can Resolve Up- and Downstream Copyright Issues

Abstract

Issues over the copyright of Large Language Models (LLMs) have emerged on two fronts: using copyrighted Intellectual Property (IP) in training data, and the ownership of generated content from LLMs. We propose adopting an opt-in system for IP owners with fair compensation determined by tagging metadata. We first suggest the development of new, multimodal approaches for calculating substantial similarity within generated derivative works by using tags for both content and style. From here, compensation and attribution can be calculated and determined, allowing for a generated work to be licensed and copyrighted while providing a financial incentive to opt-in. This system can allow for the ethical usage of IP and resolve copyright disputes over generated content.