Clinical Privacy Index The Stack Bio-Stream About Research Contact

Not automatically — and that nuance is the most important fact in European health-data strategy right now. The European Data Protection Board's draft Guidelines 02/2026 on Anonymisation (published 7 July 2026, public consultation open until 30 October 2026) treat synthetic data as anonymous only where identification of individuals is effectively precluded. Simply calling data "synthetic" earns nothing. What earns anonymity is demonstrating that the generation process prevents singling out, linkability, and inference — and the draft guidelines explicitly name differential privacy among the recognised anonymisation measures capable of doing so.

Why "synthetic" doesn't automatically mean "anonymous"

GDPR Recital 26 draws the line: data is personal if a natural person is identifiable by "all the means reasonably likely to be used." A generative model trained without privacy protection can memorize its training data — reproducing rare combinations of attributes that effectively single out real patients. Membership-inference research demonstrates this concretely: given a model's outputs, an attacker can sometimes determine whether a specific person was in the training set. In that case the synthetic dataset still relates to identifiable individuals, and the GDPR applies to it in full — special-category health data rules included.

What the draft guidelines look for

The EDPB applies the classic three-part test to any claimed anonymisation, synthetic data included:

A synthesis process must preclude all three, accounting for means reasonably likely to be used — including auxiliary datasets. This is exactly the property that ε-differential privacy formalizes: it bounds the influence of any single individual on the entire output distribution, so what an attacker can learn about any specific person is provably limited regardless of what else they know. That is why DP appears in the draft as a recognised measure while ad-hoc approaches must argue their case empirically, dataset by dataset.

Precision matters: these guidelines are a draft under public consultation until 30 October 2026 and may change on adoption. What is already clear is the direction: identification-risk tests are getting sharper, and provable methods are being distinguished from asserted ones.

What this means for health-data teams

  1. Ask any synthetic-data provider for the guarantee, not the adjective. "It's synthetic" is not a legal position. "(ε = 1.0, δ = 10⁻⁵) under Rényi-DP accounting, adjacency defined as add/remove one patient, parameters on the certificate" is one.
  2. Anonymisation is itself processing. Generating synthetic data from real patient data requires a lawful basis for that generation step; the freedom applies downstream, to the anonymous output.
  3. Document the assessment. The accountability principle means being able to show why identification is precluded: mechanism, parameters, and evaluation. A certificate with stated (ε, δ) and a tamper-evident audit trail is that documentation in filing-ready form.
  4. Watch the same logic spread. India's DPDP Act framework pushes health-data localization while verified anonymous outputs travel; the European Health Data Space builds secondary-use infrastructure that will need exactly this class of guarantee. Cross-border health research increasingly runs through provably anonymous derivatives while raw data stays home.

The strategic reading

For a decade, European health-data sharing has leaned on de-identification plus contractual controls — a structure that survives on the assumption that nobody looks too hard at the identification-risk question. The draft guidelines are the regulator looking hard. Organizations whose data strategy rests on "we removed the identifiers" have a fragility problem; organizations that can present a mathematical disclosure bound have an asset. The window between those two positions is where the next few years of health-data infrastructure will be decided. (Background: Synthetic Data vs. De-Identification.)

Frequently asked questions

Does the GDPR apply to synthetic health data?

It applies to the generation process (which processes real patient data and needs a lawful basis) and to any synthetic output that still permits identification. Synthetic data generated so that singling out, linkability, and inference are effectively precluded — for example under differential privacy with appropriate parameters — falls outside the GDPR as anonymous data.

Is differential privacy required for synthetic data to be anonymous under GDPR?

It is not the only path named, but the EDPB's draft guidelines list differential privacy among recognised anonymisation measures, and it is currently the only widely deployed approach that provides a provable bound rather than an empirical argument. Other approaches must demonstrate that identification is precluded case by case.

What epsilon value makes synthetic data anonymous?

Neither the GDPR nor the draft guidelines set a numeric threshold; the test is whether identification is effectively precluded given means reasonably likely to be used. Academic practice treats ε ≤ 1 as conservative; what matters legally is a documented assessment tying the chosen parameters to the identification-risk test.

Do these draft guidelines have legal force?

Draft EDPB guidelines are not yet adopted and may change after the consultation closing 30 October 2026. In practice, supervisory authorities and courts treat EDPB positions as heavily persuasive, so prudent teams align to the direction of travel now.

Working with clinical data that can't move?

Synthetix Health builds certified synthetic healthcare data infrastructure — ε-differential-privacy synthesis with a certificate and audit trail attached to every dataset.

Talk to us