The case, explained

AI and Privacy: DPA Ruling on Web Scraping and Model Training

6 min read · Updated September 2026 · Editorial oversight: Avv. Federico Papa

The Italian Data Protection Authority issued a guidance document on AI model training activities, marking a turning point for the regulation of massive datasets. According to press reports, recent developments have seen the launch of investigative inquiries into major tech operators using web scraping to fuel their algorithms. This intervention is part of a complex affair that began in 2023 with a fact-finding investigation aimed at regulating the indiscriminate collection of data available online. This article reconstructs the legal and jurisprudential evolution, culminating in a twin case that illustrates the practical challenges for businesses and digital professionals.

AI and Privacy: DPA Ruling on Web Scraping and Model Training

In brief

This article analyzes the lawfulness of web scraping for AI model training in light of the Italian DPA guidance. Unlike previous contributions focused on workplace transparency or biometrics, this study explores the balance between technological innovation and the right to privacy, focusing on the legitimate interest balancing test and the technical barriers required to ensure GDPR compliance.

  1. The fact

    According to reports from outlets such as Wired Italia and Il Giorno, the Data Protection Authority launched a fact-finding investigation in November 2023. The case originates from the practice of web scraping, namely the massive extraction of data from public and private websites to train generative AI models. The guidance document is part of an administrative oversight phase involving major international players, aimed at verifying whether the mere availability of data online authorized its commercial reuse. The current procedural stage is that of a fact-finding inquiry: the DPA is evaluating whether various website owners failed to implement adequate security measures, leaving user data vulnerable to automated collection. Simultaneously, the specific inquiry into OpenAI concluded on December 20, 2024, with a 15 million euro fine, clarifying the boundaries between data requirements for technological progress and the lack of explicit consent from data subjects. Regarding another well-known multinational tech company, there is no investigation by the Italian DPA, as complaints fall under the jurisdiction of other European authorities.

  2. The laws at play

    The regulatory core of the matter lies in EU Regulation 2016/679 (GDPR).

    1. Art. 5 GDPR imposes the principles of purpose limitation and data minimization, prohibiting processing for purposes incompatible with the original collection.
    2. Art. 6, par. 1, lett. f) GDPR defines legitimate interest as a legal basis, requiring a balancing test between the controller's interest and the subject's rights.
    3. Art. 25 GDPR introduces the principle of privacy by design, forcing site managers to integrate technical defenses against scraping.
    4. Art. 83 GDPR provides for fines up to 4% of annual turnover for established violations.
  3. What the jurisprudence says

    European jurisprudence has established that legitimate interest is not a blank check for online data collection. The CJEU has clarified that the balancing must be rigorous and cannot justify massive profiling if the user lacks simple tools to object. Similarly, EDPB opinions have reaffirmed that AI model training requires absolute transparency regarding sources and an effective opt-out mechanism. Domestically, previous rulings had already highlighted that the online availability of data does not transform it into freely usable public data. Merit jurisprudence has begun to recognize the right to damages when data, although accessible, is decontextualized and inserted into permanent datasets without the subject being able to reasonably foresee such use.

  4. Analysis drafted and verified with edit.legal

    To verify the provisions cited in this article, we used edit.legal. Test our legal AI on official sources and apply it to your own matters.

    Try edit.legal AI
  5. Lessons for professionals

    1. Always conduct a documented Legitimate Interest Assessment (LIA) before launching automated data collection campaigns.
    2. Implement privacy by design technical measures, such as filtering sensitive data during the extraction phase.
    3. Verify the Terms of Service of target sites to avoid contractual violations that worsen the privacy non-compliance profile.
    4. Establish channels for data subjects to exercise their rights, including simplified opt-out mechanisms.
  6. Update and rectification note (17 September 2026)

    The previous version of this article incorrectly reported the existence of an investigation by the Italian DPA into Meta, which is actually non-existent as it is the subject of a complaint to the Irish DPC. Furthermore, the proceeding against OpenAI was presented as pending, whereas it concluded on December 20, 2024, with a 15 million euro fine. Finally, the incorrect date referring to the DPA's guidance document has been removed. The text has been corrected based on official records.

References: Regolamento (UE) 2016/679 (GDPR)D.Lgs. 196/2003 (Codice Privacy)Regolamento (UE) 2024/1689 (AI Act)Provvedimento Garante Privacy su OpenAI del 20 dicembre 2024

Avv. Federico Papa
Editorial oversight: Avv. Federico Papa·ICAMContent drafted with AI support and subject to editorial source checks. Despite these controls, inaccuracies may remain: reports and rectification requests are welcome. Report a correction

Frequently asked questions

Is it illegal to download data from the internet to train an AI?

It is not inherently illegal, but it requires a valid legal basis (such as legitimate interest) and compliance with the principles of transparency and minimization set out in the GDPR.

What does a company risk by performing wild scraping?

It risks administrative fines up to 20 million euros or 4% of turnover, in addition to an order for data deletion and damages in civil court.

How can I protect my website from scraping?

It is advisable to include anti-scraping clauses in the terms of use, correctly configure the robots.txt file, and adopt technical solutions such as rate limiting and bot flow monitoring.

Verified legal research and drafting with edit.legal

Legal research and drafting with citations checked against official databases. edit.legal is free to try, no credit card.

Try edit.legal for free