Artificial intelligence (AI) is transforming the way content is created, but how should European copyright law respond to the challenge of training AI systems in a balanced way? A new IRIS report from the European Audiovisual Observatory, Copyright and AI training, examines one of the most pressing legal questions facing Europe's creative industries: how AI systems are trained and what this means for copyright of the content we are inputting. This brand-new report has been authored by Diego de la Vega, Senior Legal Analyst in the Observatory's Department for Legal Information.
As the use of AI becomes increasingly an integral part of creative production, this new report explores the legal framework governing the use of copyrighted works to train AI models, the growing importance of text and data mining (TDM), the role of user prompts, and the relationship between AI platforms and their users.
Chapter One: AI Content Generation in Context - A Copyright Perspective
Chapter One traces the explosion in AI use since the introduction of the first popular generative AI tool in 2022. The author charts their impact across Europe in both creative and business sectors and delves into the current state of the art. This report explains the technology implications of GenAI and agentic AI systems in copyright, and why these innovations present unique copyright issues, particularly when copyright-protected works are used to train AI models.
Some sensitive sectors are identified, and the author outlines why copyright has become a central concern for both content creators and those deploying AI. The landscape is made even more complex by evolving legal frameworks, including the landmark Council of Europe Framework Convention on Artificial Intelligence and the new EU Artificial Intelligence Act, both of which aim to keep pace with technological progress.
Chapter Two: The Role of Text and Data Mining (TDM) Exceptions in AI Training
Chapter two dives into the heart of the copyright debate: text and data mining (TDM) processes which are pivotal to AI training processes. The report examines the European legal provisions that allow for TDM exceptions, their roots in the Copyright in the Digital Single Market Directive (CDSMD), and their intricate relationship with national copyright laws.
This chapter analyses the opt-out mechanisms, transparency obligations, and persistent doubts as to whether TDM exceptions fully cover the wide range of AI training activities. The author shows how European, UK, and other national approaches diverge on matters of licensing, enforcement, and the balance between rightsholders and developers. Policy alternatives across the different jurisdictions range from robust rights reservation frameworks to wide exceptions for commercial and non-commercial research.
Chapter Three: Copyright in AI Training and Prompting
Chapter three explores the technical complexity of AI training and its consequences for copyright law. It presents real-world cases like the German GEMA v OpenAI ruling and breaks down the stages of AI model training where copyright rights are engaged. The chapter also demystifies "prompting" (the instructions users give to AI systems) and discusses the legal status of prompts with regard to AI training.
Chapter Four: Platforms and Users: The Terms of Service
Chapter four analyses how AI platforms' terms of service distribute copyright liability between providers and users. The report presents examples from leading platforms such as Adobe, ChatGPT, Claude, Copilot, and Midjourney, highlighting the importance of transparency, opt-out options, and contractual provisions that impact both creators and end-users.
Chapter Five: Main Takeaways
The concluding chapter distills the findings: AI technologies are advancing at a meteoric pace, but the copyright framework in Europe tries to keep up with this. The training of AI models on copyrighted data remains contentious, with text and data mining at the centre of legal uncertainty, and prompting raising new questions about human authorship and machine-generated works. Enhanced transparency, clarity, and balance (between access to datasets and the protection of rightsholder) mark the key priorities for current discussions and future policymaking.
This is a must-read for policymakers, creators, content industry professionals, legal experts, academic researchers, and anyone navigating the intersection of technology and copyright. This report provides insights, detailed comparisons, and a critical map of ongoing debates. As Europe awaits further guidance from its courts and legislators, Copyright and AI training is an essential resource for understanding where we stand, and where we might be heading, on copyright's frontier in the age of artificial intelligence. A second part exploring the copyrightability of the output produced by AI systems will follow during the second half of 2026.
Website of the European Audiovisual Observatory
The Council of Europe's work for a responsible Artificial Intelligence