Generative AI tests limits of copyright

How copyright law responds to generative AI could have major consequences for the technology’s future, says visiting Stanford Professor Mark Lemley.

Getty-1600

AI tools can regurgitate books and other copyrighted material almost word-for-word. Does that mean information inside AI models should be deemed legally recognised ‘copies’ of the original works?

It’s an unresolved question with major implications for the legality of large language models (LLMs) such as OpenAI's ChatGPT, Anthropic's Claude, and Google's Gemini

Esteemed Stanford University Professor Mark Lemley will explore this, and other copyright challenges, at a public lecture at Waipapa Taumata Rau, University of Auckland on 14 October.

Lemley, who’s teaching AI and the Law, as well as video game law to Auckland Law School students this semester, will dive into whether training AI on copyrighted material is illegal and what this means for the future of AI. 

Stanford Professor Mark Lemley
Stanford Professor Mark Lemley is known for his studies of intellectual property and internet law.

Large language models don’t store books, music or articles like files on a computer.

Instead, they learn patterns and relationships from their training data. Those learned patterns are used to predict and generate text or other content.

Because it’s possible to extract verbatim, or near-verbatim text of certain copyrighted works from some large language models, there’s evidence, says Lemley, that the model encodes the works in some form – that the model has “memorised” those works from its training data.

He says existing copyright law provides little guidance on whether this means the model contains a copy.

“The question of training is the most important one, because if copyright law bars training, there won't be any generative AI, or at least not in the form we recognise it. But even if training is legal, AI companies will be on the hook when their models generate infringing output.”

Lemley says copyright law will likely take a ‘functional approach’ to the issue, finding that LLMs contain a copy of a particular work only if it’s straightforward to extract that work.

He says this would be “an unsatisfying result”, and that changes to copyright law may be needed.

The question is one of several copyright issues now confronting courts around the world, with more than 100 lawsuits involving generative AI pending and more expected.

How those disputes are resolved could determine whether and how AI models can legally be trained, as well as how copyright law applies when models produce material that infringes existing works.

At the lecture, Lemley will also examine how likely it is that an LLM will produce infringing output, why it happens and what companies can do about it. He’ll discuss who, if anyone, owns new works made by AI, and how different countries are approaching the rapidly evolving area.

“New Zealand is still figuring out what it wants to do about AI and copyright. The proposed amendments to the Copyright Act 1994 say, in essence, ‘let's study the problem’. Other countries are adopting a wide variety of rules, from the permissive to the extremely restrictive.”

Lemley will also draw on computer science research into how generative AI models behave, an area he says remains surprisingly poorly understood despite its importance to resolving some of the questions now before the courts. 

Professor Mark Lemley is the world’s most-cited intellectual property law scholar and the third most-cited legal scholar of all time. He is the William H. Neukom Professor of Law at Stanford Law School and Director of the Stanford Program in Law, Science and Technology.

The New Zealand Centre for Intellectual Property public lecture, Copyright Law and Generative AI, is at the University of Auckland on 14 October. The lecture is supported by the Dean's Distinguished Speaker Fund, Faculty of Law, University of Auckland.

Media contact:

Sophie Boladeras, media adviser
M: 022 4600 388
E: sophie.boladeras@auckland.ac.nz