Identity graphs and audience data, explained

Synthetic Sample Expansion for Hard-to-Reach Healthcare Audiences

What synthetic sample expansion is, how it is used for hard-to-reach healthcare research audiences, how to validate it, and where it falls short.

Christian Guerrero Published 3 min read Part 9 of 10

The short answer

Synthetic sample expansion uses statistical or AI models trained on real survey responses to generate additional simulated respondents, so analysts can explore small subgroups that are hard to recruit, such as rare-disease specialists. It can help with early exploration and planning, but it cannot add information the real sample does not contain. Results should be validated against held-out real data and never presented as if they came from real respondents.

Rare-disease specialists, pediatric subspecialists, and small patient groups are hard to recruit for research. Synthetic sample expansion offers a way to stretch small samples. It is a useful tool with clear limits, and it is easy to misuse.

How it works

  1. Collect a real sample, even if small.
  2. Train a model on the responses and respondent characteristics.
  3. Generate simulated respondents that follow the patterns the model learned.
  4. Analyze the combined or synthetic data, usually with weighting.

Methods range from classical statistical approaches to large language model-based "synthetic personas." The synthetic audiences article covers related media uses.

What it can and cannot do

Can Cannot
Smooth estimates for small subgroups Add new information beyond the real sample
Help design follow-up research Replace real respondents for decisions
Explore "what if" scenarios Correct bias in the original sample
Speed early planning Provide confident statistics for tiny groups

If your real sample has 12 rare-disease specialists, generating 200 synthetic ones does not give you 212 opinions. It gives you 12 opinions, modeled.

Validation

Ask any provider:

  1. How do you validate synthetic results? Against held-out real responses?
  2. How close were predictions to real data in past studies, by question type?
  3. How do you handle subgroups with very few real respondents?
  4. How is uncertainty reported?

A good provider can show cases where synthetic results disagreed with real data and explain why.

Bias risk

Models reproduce and can amplify biases in training data. If the real sample overrepresents academic specialists, synthetic expansion will too. The AI audience model bias audit article explains how to check.

Appropriate uses

  • Early hypothesis generation before a full study.
  • Testing questionnaire design.
  • Planning sample sizes for real research.
  • Scenario exploration, clearly labeled.

Inappropriate uses

  • Presenting synthetic findings as real HCP or patient views.
  • Using synthetic data as evidence in promotional claims.
  • Making major budget decisions on synthetic subgroups alone.

Common mistakes

  • Reporting synthetic sample size as real sample size.
  • Not disclosing that results are synthetic.
  • Skipping validation against real data.

Practical takeaway

If a provider offers synthetic expansion, ask them to predict answers for a subgroup you already have real data on, then compare. Their accuracy on known data is the best guide to how much to trust them on unknown data.

Frequently asked questions

What is synthetic sample expansion?

Creating simulated survey respondents from a model trained on real responses, to increase the apparent size of small subgroups for analysis.

Is synthetic sample accurate?

It reproduces patterns in the training data. It can be useful for exploration, but it cannot discover views the real sample did not capture, and errors in the model carry through.

Should synthetic respondents be used for regulatory or medical decisions?

No. Use them, if at all, for early exploration and planning, clearly labeled, with real data for decisions.

Sources

External guidance and platform documentation change. Links were current at publication; check them again before relying on them for a decision.

Editorial note. Analysis and frameworks are the author's own and do not represent Acxiom or any current or former employer, client, or named platform. Examples labeled hypothetical or illustrative are not results from real campaigns. Nothing here is legal, regulatory, or medical advice.

Working through this decision on a real plan?

I work on health and pharma data, identity, and activation, after five years running HCP and DTC programmatic agency-side. Happy to talk through how this applies to your situation.