Article content
In briefShow moreShow lessThe EDPB published new guidelines on web scraping for generative AI and adopted final blockchain guidelines.
- The EDPB published new guidelines on web scraping for generative AI and adopted final blockchain guidelines.
- The guidance shows that publicly accessible does not mean freely reusable for any purpose.
- EDPB guidance is relevant in Norway through the EEA GDPR framework. It is not a GDPR amendment and does not mean every scraping activity is lawful or unlawful.
What happened
The EDPB published new guidelines on web scraping for generative AI and adopted final blockchain guidelines. Scraping personal data requires a basis under Article 6, while special-category data also requires an Article 9 exception.
The guidance shows that publicly accessible does not mean freely reusable for any purpose. Source, scale, expectations, transparency, objection routes and safeguards against re-identification must be assessed together.
Legal status in Norway
EDPB guidance is relevant in Norway through the EEA GDPR framework. It is not a GDPR amendment and does not mean every scraping activity is lawful or unlawful.
What the sources clarify
For special-category data, an Article 6 legal basis is insufficient; the organisation also needs an Article 9 exception. The EDPB recommended reliable sources, timestamps and validation to support accuracy, together with minimisation and technical measures limiting collection. Informing every person may sometimes be impossible or require disproportionate effort, but that does not automatically remove transparency duties.
Documentation should show the alternatives considered, how objections are handled and whether a model can reproduce identifiable training data. Public availability alone does not settle purpose compatibility or individuals' reasonable expectations.
Data mapping should trace material from source and collection date through filtering, training and deletion. The organisation should record objections and other site signals, while also assessing expectations and necessity. Memorisation and reproduction tests need realistic prompts and a remediation path when a model can reveal identifiable information.
Practical implications
Before approving a training dataset, the data owner should show source distribution, objections, special categories, filtering results and reproduction tests. Where provenance cannot be traced or objections cannot be handled, material should be isolated or removed instead of remaining an undocumented part of the model.
Sources
European Data Protection Board: “EDPB guidance on anonymisation and web scraping for generative AI,” 8 July 2026.
European Data Protection Board: “EDPB Guidelines 02/2025 on blockchain technologies,” 14 April 2025.
For discussion
Which decision should we reassess first when the legal basis changes?








