Clinical Privacy Index The Stack Bio-Stream About Research Contact

Under the European Health Data Space, a Health Data Access Body may release anonymized or synthetic outputs from a secure processing environment, but the Regulation does not define what makes such a release adequate, and that gap will be filled between now and 26 March 2029 by guidance rather than by law. Regulation (EU) 2025/327 does not use the word "synthetic" at all. It requires that whatever leaves a secure processing environment be non-personal, and it leaves the test for "non-personal" to the GDPR, to the access body, and to whatever guidance exists on the day. This note proposes five criteria an access body could apply now: a stated purpose grade, a disclosed and composable privacy budget, an attack-tested privacy claim, a provenance record a reviewer can recompute, and printed contraindications. Each can be checked by a body that does not build generative models and does not want to.

What does the EHDS actually say about anonymized and synthetic outputs?

Regulation (EU) 2025/327 was published in the Official Journal on 5 March 2025 and entered into force on 26 March 2025. It applies in general from 26 March 2027. Chapter IV, which governs secondary use, applies from 26 March 2029, and certain data categories listed in Article 51(1), among them genetic, epigenomic and genomic data, clinical-trial data and research-cohort data, follow on 26 March 2031 (Article 105).

The secondary-use chapter turns on three things. A Health Data Access Body, designated by each Member State under Article 55, decides who gets access to what. A data permit (Article 68) is the administrative decision that grants access for a stated purpose. A secure processing environment (Article 73) is the only place a permitted user may work with the data. Natural persons may opt out of secondary use under Article 71.

On outputs the Regulation is short but specific. Article 66(2) requires access bodies to provide data "in an anonymised format" where the purpose can be achieved with it; pseudonymized access is the exception and must be justified. Article 73(2) requires the access body to review every download request "to ensure that health data users are only able to download non-personal electronic health data, including electronic health data in an anonymised statistical format". Article 61(4) requires that "the results or output of secondary use shall contain only anonymous data". Article 69 allows a health data request that returns an answer only in "an anonymised statistical format", with no access to the underlying records at all.

That is the whole of it. A synthetic dataset generated inside a secure processing environment is, legally, one more candidate output: it may leave only if it is non-personal, and the access body is the one that has to say so. Recital 92 adds the warning that even "state-of-the-art anonymisation techniques" leave "a residual risk that the capacity to re-identify could be or become available, beyond the means reasonably likely to be used". The Regulation therefore hands the access body a duty (check the output) and a standard (non-personal), but no method.

Why is "anonymous" not a safe word for synthetic data?

Because a generated dataset can carry its source with it. A model trained on patient records can memorize rare records and reproduce them: Carlini and colleagues extracted near-verbatim training examples from language models in 2019 and from image diffusion models in 2023. Stadler, Oprisanu and Troncoso showed at USENIX Security 2022 that synthetic data without a formal guarantee gives no consistent protection against inference and can leak more about outliers than the original data would. Ganev's 2024 analysis constructed datasets that pass every common similarity-based privacy test while containing exact copies of real records. A similarity score is a property of one sample from the model. It says nothing about what the next sample, or a targeted query, will reveal.

The guidance texts have caught up with the literature. The TEHDAS2 joint action's guideline for access bodies on data minimisation, pseudonymisation, anonymisation and synthetic data (deliverable D7.2, 24 March 2026) states that "synthetic data is not necessarily anonymised", that it "may fall under GDPR if individuals can still be re-identified with reasonable effort", and that "it is necessary to demonstrate the resistance to re-identification". The same guideline notes that assessments based on inference attacks "are more informative and should be prioritised over methods relying solely on record-level similarity metrics". The EDPB's Guidelines 02/2026 on anonymisation, adopted for public consultation on 7 July 2026 with comments open until 30 October 2026, apply the singling-out, linkability and inference tests to synthetic datasets and warn that inferences drawn by querying models or synthetic data can breach the inference criterion. Those guidelines are a draft and may change. The direction of travel is not in doubt: under GDPR Recital 26, anonymity is judged against "all the means reasonably likely to be used", and for a generated dataset those means include querying the generator. (For the GDPR analysis in full, see Is Synthetic Health Data Anonymous Under GDPR?)

What this rules out is a release rule of the form "synthetic, therefore anonymous". What it does not yet supply is the positive rule: what the data holder must show, and what the access body must check, for a synthetic release to be adequate. "Adequate" is used here in its ordinary sense, sufficient for the access body to discharge its Article 73(2) duty with a reason it could write down.

Who is writing the missing definition?

Three tracks, none finished.

The TEHDAS2 joint action, funded to help Member States implement the EHDS, published its access-body guideline in March 2026 after a public consultation. Its final section lists what remains open: "quantitative privacy criteria and recommended parameter ranges for anonymised and synthetic data", and "a standardised, machine-readable metadata format for documenting anonymised and synthetic datasets". It reports that access bodies surveyed by the joint action used k-anonymity thresholds ranging from 3 to 100, and concludes that "it is unlikely that single fixed values for privacy parameters", including ε in differential privacy, "can be defined to cover all use cases". That is a joint action stating, in its own deliverable, that the criterion does not yet exist.

The IHI-funded FORTIFY project (grant agreement 101251225, May 2026 to April 2029) works the intellectual-property side of the same problem. The call it answered asks for "mechanisms and technologies for IPR-aware data manipulation, including reviewing best practices in anonymisation / pseudonymisation techniques and synthetic data generation", and for frameworks that may include "a classification of data into categories depending on IP sensitivity". Its public objective includes utility benchmarks and validation across ten use cases that simulate EHDS workflows.

National designation is the third track, and the slowest. Article 55 requires each Member State to designate one or more access bodies. In Sweden, the government issued preparatory assignments to four agencies in May 2026, with reports due on 1 November 2026, and on 10 September 2026 tasked two of them with proposing how a secure processing environment under the EHDS should be built, reporting by 31 March 2027; no formal designation has been made. A body that does not yet exist cannot publish criteria, which is why data holders should not wait for it.

A fourth input is the research agenda itself. van Drumpt and colleagues (Frontiers in Digital Health, 2025), from 16 expert interviews, map the risks of EHDS secondary use to privacy-enhancing technologies and close with open research questions rather than settled practice. The professional consensus is that the method is unsettled. The legal deadline is fixed.

Which five criteria could an access body apply now?

The proposal below is written for a reviewer with a statistician's training and no access to the generator. Each criterion has a form the data holder submits and a check the access body performs. None requires the access body to trust the holder's model. The criteria extend the fitness-for-purpose frame set out in Fit for Which Purpose? to the specific decision an access body has to make.

CriterionWhat the data holder showsHow the access body checks
1. A stated purpose gradeThe uses the release was evaluated for (feasibility counts, analytics test data, model pre-training, hypothesis generation), with the utility measures and results per useCompares the evaluated uses with the purpose in the data permit (Article 68) and the minimisation duty in Article 66; a release whose evaluated uses do not cover the permitted one is not adequate for that permit
2. A disclosed and composable privacy budgetThe (ε, δ) under which the data were generated and the accounting method, so that the cost of this release can be added to earlier releases from the same sourceRecords the budget against the source dataset in its catalogue and refuses a release that would exceed the ceiling the body has set for that source. Composition across releases is a theorem, not a policy choice; Dwork and Roth (2014) give the bounds
3. An attack-tested privacy claimMembership-inference and linkage results against a holdout, with the attacker model stated, presented alongside the formal bound rather than instead of itReads the attack results as confirmation of the budget; treats a similarity score offered alone as insufficient, following the TEHDAS2 guideline's own ordering
4. A provenance record a reviewer can recomputeA tamper-evident, hash-chained record of the source, the code-system versions, the generation run and the budget spent, linked so that any alteration is detectableRecomputes the record's integrity check; a record that cannot be recomputed is treated as absent. This is the artifact that extends the logging duty in Article 73(1)(e) from the session to the release itself
5. Printed contraindicationsThe uses the release was not evaluated for, or failed: rare-event detection, safety signals, and any subgroup below a stated count, stated on the certificateChecks that the permit's purpose is not on the contraindicated list; publishes the contraindications with the dataset description under Article 77

Two features of the list matter. First, criteria 2 and 5 are the only ones that can be stated before any attack is run; they are properties of the process, not of one sample. Second, every criterion is a document. An access body can apply the list with a checklist and a hash function, which is the standard the Regulation implies when it makes the access body, not the holder, responsible for the download decision. The privacy-budget primer covers what (ε, δ) states and why budgets add up across releases.

What does "accountable" mean here?

The question of whether a health-data policy can be evaluated is older than the EHDS. Faxvaag and colleagues compared the national digital-health policies of the five Nordic countries in 2024 and rated Finland and Iceland as having the most accountable policies, "meaning that their policy documents are the most transparent as to how they arrived at the conclusions and how they are to evaluate the achievements". The measure of accountability was not ambition. It was whether the document said how success would be checked.

Applied to synthetic releases, the same test reads: does the access body's rule state, in advance, what it will measure and what result will fail? A rule that says "synthetic data must be anonymous" is not accountable in this sense, because two competent reviewers can disagree about a release and neither can be shown wrong. A rule that says "a release carries a budget no greater than a stated ceiling for this source, attack results against a holdout, a recomputable provenance record and printed contraindications" is accountable, because a released dataset that lacks any of them is a documented failure, and the body that released it can be asked why.

The practical argument for Sweden, and for any Member State whose access body is not yet designated, is that the choice is not between criteria and no criteria. It is between criteria the body writes and publishes before 2029 and criteria it inherits from guidance written elsewhere, for other data. The Nordic comparison suggests which of those two positions the record rewards.

What should data holders prepare in 2027 and 2028?

Nothing in the list above requires a designated body to exist. A hospital, registry or research cohort that expects to serve data permits from 2029 can assemble the evidence pack now:

Holders that do this will find that the evidence pack also answers the data-permit application questions in Article 67 and the dataset-description duty in Article 77, because the same facts are being asked for. Holders that wait will be asked for the same facts by a body under time pressure, with no agreed format, in 2029. How this evidence pack differs from a de-identification report is set out in Synthetic Data vs. De-Identification; definitions are in What Is Synthetic Health Data?

Limits of this proposal. No Health Data Access Body has adopted the five criteria above; they are a proposal, published to be argued with. A release that carries a formal privacy guarantee still needs a per-release utility check, because a budget bounds what the data can reveal, not whether the data are useful for the permitted purpose. Synthetix Health's engine is architected to satisfy each criterion by issuing a certificate that states the purpose grades, the privacy budget, attack results, a hash-chained audit trail and contraindications; the underlying method is the subject of published patent application 202641026065 and PCT/IN2026/051623. Our own evaluation, on public data, reports a composite marginal concordance of 0.91 and 7 of 7 referential-integrity checks passed at ε = 1.0. There is no EU data-holder deployment yet, and no third-party evaluation has been published.

Frequently asked questions

Does the EHDS allow synthetic data?

The Regulation neither names nor forbids it. Article 73(2) lets a data user download only non-personal data from a secure processing environment, and Article 61(4) requires that results of secondary use contain only anonymous data. A synthetic dataset may leave if the Health Data Access Body is satisfied it is non-personal; the Regulation does not say how that is to be shown.

When does EHDS secondary use apply?

Regulation (EU) 2025/327 entered into force on 26 March 2025 and applies in general from 26 March 2027. Chapter IV, on secondary use, applies from 26 March 2029. Certain data categories, including genetic, epigenomic and genomic data and clinical-trial data, follow from 26 March 2031.

What is a Health Data Access Body?

A public body designated by each Member State under Article 55 to decide on applications for secondary use of electronic health data, issue data permits, provide a secure processing environment, and check what leaves it. Sweden's designation is pending; preparatory assignments were issued in May 2026 with reports due 1 November 2026.

Is synthetic data anonymous under GDPR?

Not automatically. Anonymity is judged against all means reasonably likely to be used, and for generated data those means include querying the generator. The TEHDAS2 access-body guideline states that synthetic data is not necessarily anonymised, and the EDPB's draft Guidelines 02/2026 apply the singling-out, linkability and inference tests to synthetic datasets. The draft is in consultation until 30 October 2026.

What is a secure processing environment?

The controlled environment, defined in Article 73, in which a permitted user may work with electronic health data: access restricted to named persons, no uploading or removal outside the permit, identifiable logs kept for at least one year, and a review by the access body of every download request so that only non-personal data leave.

Sources & further reading
  • Regulation (EU) 2025/327 of the European Parliament and of the Council of 11 February 2025 on the European Health Data Space. OJ L, 5 March 2025. Articles 51, 55, 61, 66, 68, 69, 71, 73, 77 and 105; Recital 92. eur-lex.europa.eu/eli/reg/2025/327/oj
  • Regulation (EU) 2016/679 (General Data Protection Regulation), Recital 26. eur-lex.europa.eu/eli/reg/2016/679/oj
  • European Data Protection Board. Guidelines 02/2026 on Anonymisation, version 1.0, adopted for public consultation 7 July 2026 (consultation open to 30 October 2026). edpb.europa.eu
  • European Data Protection Board. Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models, adopted 17 December 2024. edpb.europa.eu
  • TEHDAS2 Joint Action. D7.2 Guideline for Health Data Access Bodies on data minimisation, pseudonymisation, anonymisation and synthetic data. 24 March 2026. tehdas.eu (PDF)
  • Innovative Health Initiative. IHI Call 10 call text, Topic 2: Enabling and safeguarding innovation in secondary use of health data in the European Health Data Space. ihi.europa.eu/apply-funding/ihi-call-10
  • FORTIFY: Framework for Optimized Regulation, Trade Secrets, and Intellectual Property in a Federated European Health Data Space. IHI grant agreement 101251225, 1 May 2026 to 30 April 2029. CORDIS · fortify-ehds.eu
  • Regeringskansliet. Uppdrag till Socialstyrelsen och E-hälsomyndigheten att föreslå en lösning för säkra behandlingsmiljöer enligt EHDS (S2026/01715), 10 September 2026. regeringen.se
  • Faxvaag A, Reponen J, Hardardottir GA, Vehko T, Viitanen J, Eriksen J, Koch S, Nøhr C. Towards accountable e-health policies in the Nordic countries. Stud Health Technol Inform. 2024;316:339–343. doi:10.3233/SHTI240413
  • van Drumpt S, Chawla K, Barbereau T, Spagnuelo D, van de Burgwal L. Secondary use under the European Health Data Space: setting the scene and towards a research agenda on privacy-enhancing technologies. Front Digit Health. 2025;7:1602101. doi:10.3389/fdgth.2025.1602101
  • Government Offices of Sweden. Fyra myndigheter förbereder Sverige för EU:s nya hälsodataförordning EHDS. Press release, 5 May 2026. regeringen.se
  • Stadler T, Oprisanu B, Troncoso C. Synthetic data — anonymisation groundhog day. USENIX Security 2022. usenix.org
  • Ganev G. Synthetic data, similarity-based privacy metrics, and regulatory (non-)compliance. GenLaw workshop, ICML 2024. arXiv:2407.16929
  • Carlini N, Liu C, Erlingsson Ú, Kos J, Song D. The Secret Sharer: evaluating and testing unintended memorization in neural networks. USENIX Security 2019. arXiv:1802.08232
  • Carlini N, Hayes J, Nasr M, et al. Extracting training data from diffusion models. USENIX Security 2023. arXiv:2301.13188
  • Rocher L, Hendrickx JM, de Montjoye Y-A. Estimating the success of re-identifications in incomplete datasets using generative models. Nat Commun. 2019;10:3069. doi:10.1038/s41467-019-10933-3
  • Dwork C, Roth A. The algorithmic foundations of differential privacy. Found Trends Theor Comput Sci. 2014;9(3–4):211–407. doi:10.1561/0400000042

Preparing a data holder, or an access body, for 2029?

Synthetix Health issues a certificate with every dataset that states the purpose grades, the privacy budget, attack results, a hash-chained audit trail and contraindications. We publish our limits alongside our numbers, and we would rather be argued with than agreed with.

Talk to us