← Blog
European Union22 July 2026DILAIG

GDPR and generative AI, what the EDPB's new guidance on anonymisation and web scraping changes

On 8 July 2026, the EDPB adopted two major texts for generative AI actors, a clarified framework on data anonymisation and guidelines on web scraping. A useful reminder for organisations already working on AI Act compliance, since these two regulations move forward in parallel and overlap on several points.

On 8 July 2026, the European Data Protection Board (EDPB) adopted two new texts that directly concern organisations developing, training or deploying generative AI systems. These guidelines are open for public consultation until 30 October 2026, but their content already gives a clear sense of where European regulation is heading.

A clearer framework for anonymisation

The first question many organisations handling large volumes of data ask themselves is simple. Is a piece of data still personal data once it has been processed, aggregated or transformed?

The EDPB answers by drawing in particular on the Court of Justice of the European Union ruling of 4 September 2025 (case C 413/23 P, EDPS v SRB). Data is considered anonymous if it no longer relates to an identified or identifiable natural person, based on means reasonably likely to be used to distinguish that person in a given context.

To assess whether anonymisation has been successful, the EDPB proposes a three criteria framework.

First, the absence of record isolation for any individual entry. Second, the absence of possible linkage between different datasets. Third, the absence of inference that would allow deducing information about a person. If all three criteria are met, the data can be considered anonymous.

Two approaches are available to apply this framework. A contextual approach, which takes into account the actual capabilities of different actors who might attempt reidentification. And a simplified approach, more cautious and easier to implement, which does not distinguish between these capabilities and applies a higher default standard.

Web scraping under closer scrutiny

The second text tackles a practice that is central to training generative AI models, large scale automated data collection from the web.

The EDPB confirms that as soon as a scraping operation involves personal data processing, whether collection, storage, organisation or retrieval, the GDPR applies in full. Three principles stand out.

The purpose limitation principle, which requires clearly defining why data is being collected. The transparency principle, with a bounded tolerance when individually informing data subjects proves impossible or disproportionate. And the accuracy principle, for which the EDPB recommends collecting only from reliable sources, timestamping data and validating it before using it to train a model.

On the legal basis question, the EDPB expands on the use of legitimate interest in the specific context of scraping for AI training, building on its earlier Opinion on AI models (Opinion 28/2024).

The most delicate point concerns special categories of data, those relating to health, political opinions, sexual orientation or other sensitive attributes. Processing them remains prohibited in principle, unless an exception under Article 9 of the GDPR applies. The EDPB discusses the possibility of invoking incidental or residual collection of such data under certain conditions, drawing on the GC and Others ruling (case C 136/17), while stressing that there is no general exemption and each situation must be assessed individually.

Why this also matters for your AI Act compliance

These guidelines address the GDPR, not the AI Act. But in practice, the two regulations regularly intersect for organisations developing or deploying AI systems.

The technical documentation required under Annex IV of the AI Act, or the fundamental rights impact assessment (FRIA) for high risk systems, often need to address the origin and nature of training data. If that data comes from scraping, the question of its GDPR legal basis and any processing of special categories of data becomes a direct point of attention for properly documenting these AI Act obligations.

Similarly, a solid data governance policy, one of the five documents organisations need to produce as part of their AI Act compliance, benefits from including a reflection on anonymisation, if only to distinguish what still falls under GDPR from what no longer does.

GDPR and the AI Act are not two separate workstreams. They are two compliance layers that often apply to the same data and the same systems. Our tool focuses on the AI Act, but we consistently recommend addressing both texts in a coordinated way rather than in silos.

Where does your AI Act compliance stand

Before diving into these GDPR questions, it helps to know exactly where you stand on the AI Act. Two free DILAIG tools let you check in a few minutes, no account required.

The prohibited practices detector (Article 5) tells you whether your AI system falls under a use banned by the regulation. Test the Article 5 detector

The GPAI qualifier determines whether your model falls within the specific obligations for general purpose AI models (Article 53). Test the GPAI qualifier

If your organisation deploys or provides a high risk system, the full DILAIG audit generates your compliance score out of 100 along with the five expected regulatory documents, including the FRIA and the data governance policy mentioned above. Start the full audit

22 July 2026DILAIG
All articles

Take action

Is your AI system compliant?

Free audit in 20 minutes. Detailed report, no commitment.

Start the audit →

Keep reading

Practical guides, regulatory analysis, DILAIG news.

View all articles →