AI and Copyright Law - Who Owns What a Model Creates
Generative AI systems raise copyright questions at two separate points: what goes into training, and what comes out as output. These are legally distinct issues, and courts and legislators are still working through both.
Training Data Questions
Large models are trained on vast datasets scraped from the public internet, much of which is copyrighted. The central legal debate is whether this training use qualifies as fair use (in the U.S.) or falls under a text and data mining exception (in some other jurisdictions):
Argument for fair use: training is transformative, output is not a copy of any single work
Argument against fair use: models are trained on unlicensed copies, and can sometimes reproduce
training data closely, especially for memorized or duplicated content
Multiple ongoing lawsuits against AI companies center on this question, and outcomes differ by jurisdiction—so far without a single settled global standard.
Output Ownership
Most copyright offices, including the U.S. Copyright Office, have taken the position that works generated entirely by AI without meaningful human creative input are not eligible for copyright protection, because copyright law traditionally requires human authorship. This creates an unusual gap: a fully AI-generated image may be legally unprotectable, while a human-edited or human-curated combination of AI outputs may qualify, depending on how much human creative choice was involved.
Style and Likeness
A separate concern is training or prompting a model to imitate a specific living artist’s style or a person’s likeness. Style itself is generally not copyrightable, but:
- Reproducing a specific existing work too closely can still infringe, regardless of how it was generated.
- Right-of-publicity and trademark laws (not copyright) may apply to voice, likeness, or persona.
- Many platforms now impose their own policy restrictions distinct from what the law strictly requires.
Practical Guidance
Organizations using generative AI commercially generally should track provenance of training data where possible, review outputs for close similarity to known copyrighted works before publishing, and document the human creative contribution to any AI-assisted output to support a copyright claim. Because case law is developing quickly and differs by country, this remains one of the most legally uncertain areas of applied AI.