Licensing Training Data and Attributing Copyright of Derivative Content From Large Language Models Can Resolve Up- and Downstream Copyright Issues

Abstract
Issues over the copyright of Large Language Models (LLMs) have emerged on two fronts: using copyrighted Intellectual Property (IP) in training data, and the ownership of generated content from LLMs. We propose adopting an opt-in system for IP owners with fair compensation determined by tagging metadata. We first suggest the development of new, multimodal approaches for calculating substantial similarity within generated derivative works by using tags for both content and style. From here, compensation and attribution can be calculated and determined, allowing for a generated work to be licensed and copyrighted while providing a financial incentive to opt-in. This system can allow for the ethical usage of IP and resolve copyright disputes over generated content.
See also
- Weakening the Voting Rights Act reduces minority representation and electoral competitionPreprint, 2026
- Governance at a Crossroads: Artificial Intelligence and the Future of Innovation in AmericaSSRN, 2025
- Pheromone-based Learning of Optimal Reasoning PathsarXiv, 2025
- MINDSTORES: Memory-Informed Neural Decision Synthesis for Task-Oriented Reinforcement in Embodied SystemsICLR Workshop on Reasoning and Planning for LLMs; arXiv, 2025
- Insights into Flexible Bioinspired Fins for Unmanned Underwater Vehicle Systems through Deep LearningBiomimetics; NeurIPS Workshop on Machine Learning and the Physical Sciences, 2024