Clinical Privacy Index The Stack Bio-Stream About Research Contact

Synthetic Medicare claims data is artificially generated beneficiary, enrollment, and claims records that look and behave like real Medicare data — same tables, same code systems, same utilization patterns — without containing any actual beneficiary's information. It exists in two generations. The first is CMS's own static public files, built a decade ago to prove the concept. The second is on-demand, certified synthesis with a formal privacy guarantee attached. Understanding the gap between the two is the fastest way to understand where health-data infrastructure is heading.

Why CMS built SynPUF — and what it proved

In 2013, CMS released the Data Entrepreneurs' Synthetic Public Use File (DE-SynPUF): synthetic versions of 2008–2010 Medicare claims for roughly 2.3 million synthetic beneficiaries, spanning beneficiary summaries, inpatient, outpatient, carrier, and prescription-drug claims. The stated purpose was exactly the use case analytics teams still have today: let developers build and test software against claims-shaped data "without the risk of violating beneficiary privacy," with programs designed so that code written against the SynPUF would run on CMS's real Limited Data Sets.

SynPUF's real achievement was proving demand. It has been downloaded, converted (including into the OMOP Common Data Model, mirrored on the AWS Open Data registry), and built against for over a decade — evidence that an entire industry needs realistic, linkable claims data it is allowed to touch.

Where SynPUF falls short today

LimitationConsequence
Frozen in 2008–2010ICD-9 era: no ICD-10, no modern therapies, pre-ACA utilization patterns
Fidelity deliberately limitedCMS intentionally distorted inter-variable relationships for privacy, and documents that analytic inferences from SynPUF should not be trusted
One schema, one populationMedicare FFS only; cannot be shaped to a customer's tables, vocabularies, or volumes
No formal privacy guaranteeProtection rests on the distortion process, not on a stated, verifiable bound
Never refreshedThe file is the file; there is no pipeline behind it

None of this is criticism — SynPUF was built as a proof of concept and says so. But teams still standardizing test environments on a 2010 snapshot are solving a 2026 problem with 2013 infrastructure.

The successor: certified synthesis

Modern synthesis inverts the model. Instead of one static file, a generation pipeline runs against whatever claims source the data holder has — and every output ships with evidence:

A worked example — with real numbers

To make this concrete: in August 2026 we ran the SX-Engine against the public 100,000-person OMOP conversion of DE-SynPUF itself (a deliberately safe seed — it is already synthetic and public domain, so the demonstration involves no PHI at any point). Output: a 5,000-person cohort across five OMOP CDM tables — person, visit, condition, drug, procedure — totalling 1.85 million rows, generated under a composed budget of ε = 1.0, δ = 10−5 with Rényi-DP accounting.

That is the generational difference in one paragraph: not "trust us, it's synthetic," but a dataset that arrives holding its own evidence.

What claims synthesis is honestly for

Same discipline as synthetic health data generally: excellent for test/dev/QA environments, demo data, analytics-workflow validation, and feasibility work; moderate for population-level exploration; not for pharmacovigilance, rare-disease cohorts, or regulatory submissions as primary evidence. A claims dataset with a certificate should print its contraindications on the certificate.

Frequently asked questions

Is CMS DE-SynPUF still available?

Yes — the 2008–2010 files remain downloadable from CMS, and OMOP CDM conversions are mirrored on the AWS Open Data registry. CMS has also published newer Synthea-based synthetic Medicare claims files. All remain static snapshots rather than generation pipelines.

Can I still use SynPUF for software testing?

For basic pipeline bring-up, yes. Its limits arrive quickly: ICD-9-era coding, deliberately distorted relationships that CMS itself warns against analyzing, and a single fixed schema. Teams needing current vocabularies, their own table shapes, or defensible privacy documentation outgrow it.

Does synthetic Medicare data involve real beneficiaries?

DE-SynPUF was derived by CMS from real claims through a disclosure-limitation process. Certified synthesis states the relationship formally: under ε-differential privacy, the influence of any single individual in the source data on the entire output is mathematically bounded, whatever the source.

Can synthesis produce ICD-10-era or payer-specific claims?

Yes — a synthesis pipeline preserves whatever coding and schema the source uses, so ICD-10/current-vocabulary output comes from current source data, shaped to the tables and volumes an engagement defines. That is the structural difference from any static public file.

What should a buyer ask for with any synthetic claims dataset?

Four artifacts: the stated privacy parameters (ε, δ, adjacency, accounting method), a referential-integrity report, a fidelity report against the source, and provenance/audit documentation. If a provider can't produce them, the word 'synthetic' is doing all the work.

Working with clinical data that can't move?

Synthetix Health builds certified synthetic healthcare data infrastructure — ε-differential-privacy synthesis with a certificate and audit trail attached to every dataset.

Talk to us