Copyrighted Works for AI Model Training

If AI needs copyrighted work to function, the question needs to shift from “Is it fair use?” to “Who gets paid, and how much?”

Generative AI models are trained on enormous amounts of data scraped from the internet, much of it copyrighted material, including books, music, films, images, code, journalism, and art.

The U.S. Copyright Office, in its May 2025 Pre-Publication Report on Copyright and Artificial Intelligence, confirmed what many creators already suspected:

The Legal Battlefield: Fair Use and AI Training

Under 17 U.S.C. §107, courts evaluate the fair use defense using four factors:

AI companies argue that training models is a “transformative” use — similar to how search engines index websites. Creators argue that generative models compete directly with the original works, undermining licensing markets.

Multiple lawsuits by authors, visual artists, and media companies are now testing that question in federal courts. Outcomes have lately teetered in favor of AI labs — but the question remains unsettled, and courts are split.

The lesson from the courts is clear: creators should not wait for a judicial resolution before negotiating. Licensing deals can happen now.

Why Licensing Makes Strategic Sense

The U.S. Copyright Office report suggests voluntary licensing as the path forward. Where voluntary licensing is infeasible, alternative frameworks — compulsory or extended collective licensing — may fill the gaps.

There is also a compounding strategic logic here: the more licensing markets develop for AI training data, the stronger market harm claims become. When a licensing market exists, courts are less able to view unlicensed use as fair use.

The goal is simple:

If your work helps train an AI model, you should participate in the value that accrues to it.

The Core Terms Creators Must Control

When negotiating AI training licenses, focus on five issues.

  1. Training Scope: Define whether the license covers model training, fine-tuning, retrieval systems, or synthetic data generation. Not all training uses are equal — and vague language will be exploited.

  2. Output Restrictions: Prevent the model from generating works that are substantially similar to your copyrighted works.

  3. Attribution and Metadata: Require preservation of authorship data and digital provenance. This is the foundation of accountability in AI-generated content.

  4. Revenue Participation: Training data is a production input. Treat it like one. Negotiate for a meaningful share of the value your IP helps create.

  5. Audit and Transparency Rights: If you can’t see how your IP is being used, you can’t enforce your rights. Audit provisions are non-negotiable.

The Strategic Reality

The AI industry is discovering the same truth Hollywood learned long ago: content is infrastructure.

Creators who license their IP today will shape the rules of the AI economy. Creators who don’t may simply become the training set.

Melissa Goodwin
I make things.
Previous
Previous

Streaming Giants & Live Sports

Next
Next

How to Avoid the IP Dealmaking Mistakes That Kill IP Businesses