EUIPO Report: Generative AI and Copyright in the EU

On May 12, the EUIPO (European Union Intellectual Property Office) published a report titled The development of Generative Artificial Intelligence from a Copyright perspective“.

This report was conducted by a research group from the School of Law at the University of Turin and the Nexa Center for Internet & Society at the Polytechnic University of Turin.

Specifically, the report analyzes the relationship between the development of Generative Artificial Intelligence (GenAI) and the copyright system within the European Union. The research focuses on the technical, legal, and economic aspects of using protected content for training and employing GenAI systems, highlighting in particular the implications for copyright holders and the providers of such systems.

Training GenAI Models and Input Challenges

GenAI systems operate by analyzing large quantities of data, often collected from the web through web scraping techniques. This data, frequently protected by copyright, is used in various phases of system training (such as pre-training, fine-tuning, and reinforcement learning).

European copyright law is based on the general Directive 2001/29/EC (Infosoc). More recently, Directive (EU) 2019/790 (Copyright in the Digital Single Market, CDSM) was added to adapt copyright law to technological evolution and the digital context, along with the AI Act, which governs the legal aspects related to the use of artificial intelligence.

In particular, the CDSM provides exceptions to the exclusive right of reproduction for copyright holders for Text and Data Mining (TDM) activities carried out for scientific research purposes and for any other purpose (Articles 3 and 4 of the CDSM). However, it grants copyright holders the possibility to object to such uses by expressing a reservation (opt-out), which must be clearly and machine-readably expressed.

Compliance with these reservations is mandatory for GenAI system providers under the AI Act, which also requires the publication of summaries of the data used for training to allow copyright holders to easily exercise their opt-out rights. However, the adoption of effective technical tools for expressing and applying TDM reservations remains an open challenge due to the heterogeneity of content and sectoral practices.

Content Generation and Output Transparency

The content generation process in GenAI models depends on the type of model and the content produced. Given the high costs and difficulties associated with continuous training, the report illustrates how technologies like Real-time Augmented Generation are becoming increasingly widespread. These integrate training-based content generation with real-time retrieval of updated information available online.

In any case, the AI Act mandates transparency for content generated by GenAI systems, and the report in question analyzes various solutions for identifying and flagging AI-generated content, highlighting their advantages and limitations. Specifically, tools have been developed to reduce the risk of copyright infringement, such as filters and techniques to modify or update systems after their release (known as “model unlearning” and “model editing”), also offering legal guarantees to users.

However, these tools have technical and application limitations that necessitate a more active involvement of public institutions in promoting interoperability and standardization of output transparency measures.

Economic Perspectives and Development of a Licensing Market

Recently, there has been growing interest in direct licensing models between copyright holders and GenAI providers. The development of a sustainable licensing market requires effective technical tools to easily exercise opt-out, transparent remuneration mechanisms, and reliable intermediary structures to facilitate the matching of content supply and demand.

The role of public authorities is therefore fundamental in this area, both to promote the adoption of technological tools and to ensure fair and legitimate access to content by GenAI system providers.

Conclusions

The report thus highlights the need for a balanced approach that can combine the protection of copyright with the opportunities offered by technological innovation. The complexities introduced by Generative Artificial Intelligence systems, both in the data acquisition phase (input) and content production phase (output), make the adoption of harmonized regulatory and technical solutions essential.

In this context, the EUIPO’s initiative to establish a Copyright Knowledge Centre by 2025 represents a significant first step in this direction. This office will provide copyright holders with practical information on how their content can be used by GenAI systems, the solutions available to them to protect their rights, as well as exchange platforms and tools to facilitate interaction among the various stakeholders involved.

Elena Bandinelli