On 15 April 2026 the European Data Protection Board adopted draft Guidelines 1/2026 on processing personal data for scientific research purposes and opened them for public consultation until 25 June. If you process personal data for research in Europe, this document is going to settle arguments you will be having with your DPO for the next ten years. And the examples included are just amazing! Great materials for teachers and learners on data protection for research.
Let’s continue with more awe: the submissions are public! Anyone could write in, and 132 organisations and individuals did. Universities, hospitals, biobanks, ministries, data protection authorities, business associations, NGOs, and even a handful of private citizens.
So naturally I thought: I’ll read them all.
Then I started checking the submitted files… counted the pages… 1050. One submission is 67 pages long, some are a single polite paragraph, and Finland alone sent 11 of them ( we are a small country that reeeaally likes data protection…). So yeah, I did not have the time to read 1050 pages, and most likely neither have you (but I am sure someone at EDPB has read them all even multiple times!).
So: how can I explore quickly the 132 submissions? I vibe-coded SciResGDPR, a tool to explore all 132 of them. Code is on GitHub, data is CC-BY, software is MIT, go play with it now!
Yes, I vibe coded it. But I want to be precise about what that means here, because “vibe coding” has become a synonym for “slop” and I guarantee this is not that. I knew exactly what I wanted before I opened the editor: paragraph as the unit of analysis, issue and stance stored in separate fields, no forced single label, every code clickable back to the paragraph in the original PDF so that nobody (including me) has to trust the machine. The model wrote the code after a very clear plan I gave it on what I wanted to have and no fundamental decisions were made by the tool. All code is public and written with reproducibility in mind, so that at the next consultation it might be easier to re-use it.
I posted it on LinkedIn expecting my usual three likes from the data steward/data protection crowd, and instead quite a few people got interested, which is how I ended up being invited by Rosalia Anna D’Agostino at Legal4Tech for a podcast chat about the tool and the submissions. (If you don’t follow Legal4Tech, do yourself a favour and check their linkedIn and all the past podcasts, for examples staying human in the age of AI with Sam Illingworth, on how students actually use AI and why judgement is a core skill and not a soft one; AI standardisation with Ley Muller, on how political the writing of European standards really is; making data protection authorities accountable with Mark Leiser, including the excellent point that “non-personal data” is not one thing; and GEMA v SUNO on training data, memorisation and fair remuneration).
What the tool actually does (and what it absolutely does not)
So let’s start with the pipeline: first it extracts paragraphs from the PDFs (7116 of them, about two thirds substantive, the rest were just letterheads and “yours sincerely”). Then, it asks a large language model (in this case Claude Opus) to do open coding on each paragraph with zero to five short codes, keep the issue separate from the stance (support, oppose, request clarification, propose change, concern), embed the 5095 distinct code phrasings, merge them with Leiden clustering into 565 canonical codes, cluster again at two resolutions into 17 themes and 97 subthemes.
Now the caveats, and please read these before you quote any number from the dashboard at me.
This is an index, not an actual qualitative analysis. It borrows the vocabulary of qualitative coding (codes, themes, stances) because that vocabulary is legible to the people who do this work properly. But there is no constant comparison, no second coder, no reflexivity, no saturation, no negotiated agreement. Real thematic analysis is interpretive (and human!) and this is not. What the tool does is let a human who wants to do the real work find the relevant 40 paragraphs instead of reading 7116.
The stance counts are not votes. A consultation is self-selected, and the genre rewards “propose change”. Out of 6330 codes: 2200 propose a change, 1539 raise a concern, 1312 ask for clarification, 545 support, 447 oppose. Nobody wins a consultation by having more bullet points.
The clustering is not perfect. Three separate canonical codes all called “broad consent”, because the same phrase carries opposite positions and embedding similarity cannot see that but I kept them anyway. It is an honest picture of the mess of working with unstructured data, but at the end this is just a clever index of the submissions, and the reader can just click and go back to the original and read (and code) things for themselves.
What did the tool find?
One: let’s agree to not agree on the definition of research
The definition of research present in the guidelines is addressed by 108 of the 132 submissions and it is the most polarised theme in the whole corpus. Paragraph 11 of the draft sets out six key-indicative factors: methodical and systematic approach, ethical standards, verifiability and transparency “normally following peer review”, autonomy and independence (with a PhD given as the example qualification), contributing to society’s knowledge and wellbeing, and potential to contribute to existing knowledge. Is this definition too strict? Too academic? Industry and market research might not necessarily need all the six factors and still conduct very useful research. On the other hand, some organisations are also very happy with the definition: Tampere University argues that the six factors match what Finnish case law. And then noyb and an individual submitter argue the factors are too weak, because they are substitutable and get tested once and never rechecked when an outfit quietly drifts from research to product.
My own view, as someone who teaches research ethics: research is defined by method and by integrity norms (hello ALLEA), not by your employer/country/continent. And the commercial purpose does not disqualify you from being a researcher. The reason research gets privileges under the GDPR is of course the public benefit, and the factor that carries that justification is validity and transparency of the methods. Peer review is one route to check validity, but it is not the only one.
Recital 159 from the GDPR leaves a broader context for research, why not just sticking to that? And let’s add more mess: the Digital Omnibus! The digital omnibus has also proposed a definition of scientific research into a new version of Article 4 GDPR (source) with three of the EDPB’s six factors (“scientific research” means any research which can also support innovation, such as technological development and demonstration. These actions shall contribute to existing scientific knowledge or apply existing knowledge in novel ways, be carried out with the aim of contributing to the growth of society´s general knowledge and wellbeing and adhere to ethical standards in the relevant research area. This does not exclude that the research may also aim to further a commercial interest.’). The GDPR side of the Omnibus is still in Council so let’s see what will actually happen when it comes to research.
Two: consent, and what it costs open science
The consent theme in the corpus runs 43 oppose against 14 support. That is the research community telling the EDPB, fairly loudly, that consent is the wrong default legal basis for research.
The Guidelines do open the door to broad consent and to dynamic consent, and that is genuinely welcome. But look at the shape of the document: a long, detailed chapter on consent, and a much shorter one on Article 6(1)(e). This is exactly what we raised in Aalto’s submission. In Finland most research organisations have relied on Article 6(1)(e) together with Section 4(1)(3) of the Data Protection Act for the eight years since the GDPR started applying. Promoting broad or dynamic consent without equally detailed guidance on the public-interest basis may in practice push institutions into changing legal basis, and we asked the EDPB for guidance that does not presuppose that change.
Why does this matter so much for open science? Because consent as a legal basis is not a one-off action, it is a permanent monitoring obligation attached to the dataset. You must be able to demonstrate it, honour withdrawal, act on erasure. That means keeping the link to the person for as long as the data lives, keeping contact details, accepting that people withdraw non-randomly and your cohort quietly becomes biased, and knowing that if the form turns out to have been defective you cannot simply switch basis later.
The other Aalto point is the same problem from the other end. Paragraph 87 of the draft says controllers should not knowingly delete contact details if they anticipate further research. About a dozen submissions objected, ours included, because in real projects direct identifiers are deleted early precisely because they were only needed for recruitment and scheduling, and Articles 5(1)(c), 5(1)(e) and 11 say you should not keep identifying information you no longer need. Deleting direct identifiers is a safeguard, not a compliance failure. Being told to retain contact details “just in case” inverts data minimisation in the name of transparency.
And then the case that really is open science: a public data repository holding pseudonymised data from people you can no longer contact, giving controlled access to other researchers through a Data Access Committee. That is not an edge case, it is the standard operating procedure in Europe. The European Genome-phenome Archive works exactly like this. We asked the EDPB to add an example acknowledging it, because guidance written around the single project with a live participant relationship simply does not describe how data gets reused and shared today.
Three: anonymisation and pseudonymisation
The most concern-heavy theme in the whole corpus, raised by 71 submissions. Is pseudonymised data personal data for a recipient who holds no key, after the Court’s relative approach in EDPS v SRB? Can genomic data be anonymised at all, when a whole genome is unique by construction and something like 30 to 80 independent SNPs already single you out? Do secure processing environments and federated analysis count as Article 89 safeguards, or are they just nice engineering?
My answer, from the side of the fence where I actually run the machines: you cannot win this argument on the data, only on what the recipient can do with it. No download, analysis inside the environment, output checking, and a legal rather than merely contractual bar on re-identification attempts.
Which is conveniently also the subject of the EDPB’s other consultation, the one that is open right now: Guidelines 02/2026 on Anonymisation, feedback until 30 October 2026. If you work with sensitive data, this is the one to write to. It takes an afternoon, and your afternoon becomes part of a public record that someone like me will eventually index.
And on Friday 18/September
I dir a recording with Rosalia last Friday (18/09) and… I think it went really well! I am not going to spoil it here, so you will have to listen when it comes out.
One thing I do want to say: I hope I do not come across as a ranting open science fanatic willing to sacrifice data protection in the name of Science. I am not. I spend most of my time building the boring machinery that keeps sensitive data safe, and I teach doctoral students why it matters, and I also practice what I preach in all the projects I join.
But the balance between openness of research and protection of “secrets” is never black and white, and treating it as if it were is how we end up with the worst of both worlds: research that cannot be reproduced and people whose data is no better protected for it. With the right technical and organisational measures we can have reproducible science built on secondary use of personal data AND compliant data protection. That is the work in progress right now, but I am optimist that we will get there!
Anyway, explore the 132 submissions: Go read them, and find what actually matters to you.