Artificial intelligence represents one of the most revolutionary innovations of the digital age, offering extraordinary opportunities in numerous sectors, from healthcare to cybersecurity, to communication and industrial automation. However, its development and use raise critical issues regarding the protection of personal data.
To respond to these challenges, the European Data Protection Board (EDPB) has published Opinion 28/2024, at the request of the Irish Data Protection Authority, with the aim of clarifying some fundamental aspects of the processing of personal data in the context of artificial intelligence.
The opinion, adopted on December 17, 2024, addresses four main issues: the possibility of considering AI models anonymous, the use of legitimate interest as a legal basis for data processing, the consequences of unlawful processing in the development phase of an AI model, and possible measures to reduce risks to the privacy of data subjects.
Anonymization and AI Models
A central theme analyzed by the EDPB concerns the possibility that an artificial intelligence model may be considered anonymous and, therefore, not fall within the scope of application of the GDPR. The Committee emphasizes that a model trained with personal data cannot automatically be considered anonymous, as it may still contain information attributable to specific individuals. In order for a model to be effectively defined as anonymous, it must be demonstrated that the risk of extracting personal data is insignificant and that there is no possibility of identifying, even indirectly, individuals through queries or attacks on the system.
The EDPB highlights how this assessment must be conducted on a case-by-case basis, analyzing various factors, including the anonymization techniques adopted, the risk of re-identification through cyber attacks, and the context of use of the model. For example, there are advanced attack techniques, such as membership inference attacks and model inversion attacks, which can allow us to trace the original data used for training.
Therefore, if a model cannot be considered anonymous, it must comply with the provisions of the GDPR regarding the protection of personal data.
Legitimate Interest as a Legal Basis
Another central aspect analyzed by the Committee concerns the use of legitimate interest as a legal basis for the processing of personal data in the development and implementation phases of an AI model. According to the GDPR, processing is lawful if it pursues a legitimate interest, is necessary to achieve this interest, and the fundamental rights and freedoms of the data subjects do not prevail.
The EDPB clarifies that the legitimate interest must be well defined, real, and not purely hypothetical. In the field of artificial intelligence, legitimate interests may be considered, for example, the development of virtual assistants to facilitate interaction with users or the creation of systems to improve cybersecurity through the detection of fraudulent activities. However, it is necessary to demonstrate that the data processing is actually necessary to achieve these objectives and that there are no less invasive alternatives.
A particularly important element is the assessment of the impact of the processing on the rights of the data subjects. The Committee emphasizes how the complexity of the technologies used in AI models can make it difficult for users to understand the use of their data. Therefore, it is essential to consider their reasonable expectations: if a user, for example, has published some information online without the intention that it be used to train an AI model, the processing may not be justifiable on the basis of legitimate interest.
The Impact of Unlawful Processing in the Development Phase
A further issue addressed by the EDPB concerns the consequences of unlawful processing of personal data in the development phase of an artificial intelligence model. The Committee analyzes various scenarios to assess how the illegitimacy of the initial phase can affect the subsequent operation of the model.
In the first scenario, the model retains personal data and is used by the same data controller. In this case, it is necessary to verify whether the development and implementation represent distinct phases of the processing and whether the absence of a valid legal basis in the first phase makes the subsequent use unlawful.
In the second scenario, the model is transferred to another data controller who uses it without having taken part in the development phase. Here, the EDPB clarifies that the new data controller has the obligation to verify whether the model was developed in compliance with the GDPR, ensuring that the processing of personal data is lawful.
Finally, the Committee analyzes the case in which a model is anonymized after being developed unlawfully. If the anonymization process is effective and the model no longer processes personal data, the GDPR does not apply. However, in the event that the model is subsequently used to process new personal data, these must be managed in full compliance with the legislation.
Mitigation Measures and Obligations for Data Controllers
To reduce the risks to the privacy of data subjects, the EDPB recommends a series of mitigation measures that data controllers should adopt. Among these, particular attention should be paid to limiting the collection of personal data during the model training phase, favoring the use of anonymized or synthetic data where possible.
It is essential to apply protection techniques such as pseudonymization and advanced anonymization, as well as conducting regular security tests to prevent inversion and data extraction attacks.
A crucial aspect concerns transparency: data controllers must provide clear information to data subjects on how their data is processed, guaranteeing them the possibility of exercising their rights, such as the right of access, rectification, and opposition to the processing.
Conclusions
The EDPB opinion provides a clear framework on the challenges that artificial intelligence poses in the field of personal data protection and reiterates the importance of ensuring compliance with the GDPR throughout the life cycle of AI models.
Among the main points that emerged, it is highlighted that AI models trained with personal data cannot be considered automatically anonymous and must comply with data protection regulations. Furthermore, the use of legitimate interest as a legal basis must be carefully assessed, taking into account the expectations of the data subjects and the principle of data minimization.
The unlawful processing of data in the development phase can influence the legitimacy of the subsequent use of the model, making mitigation measures essential to reduce privacy risks. With the evolution of technology and regulations, it is crucial that companies and developers adopt a proactive approach to data protection, ensuring a balance between innovation and the fundamental rights of individuals.
Teresa Franza