Synthetic Sample Expansion for Hard-to-Reach Healthcare Audiences
What synthetic sample expansion is, how it is used for hard-to-reach healthcare research audiences, how to validate it, and where it falls short.
The short answer
Synthetic sample expansion uses statistical or AI models trained on real survey responses to generate additional simulated respondents, so analysts can explore small subgroups that are hard to recruit, such as rare-disease specialists. It can help with early exploration and planning, but it cannot add information the real sample does not contain. Results should be validated against held-out real data and never presented as if they came from real respondents.
Rare-disease specialists, pediatric subspecialists, and small patient groups are hard to recruit for research. Synthetic sample expansion offers a way to stretch small samples. It is a useful tool with clear limits, and it is easy to misuse.
How it works
- Collect a real sample, even if small.
- Train a model on the responses and respondent characteristics.
- Generate simulated respondents that follow the patterns the model learned.
- Analyze the combined or synthetic data, usually with weighting.
Methods range from classical statistical approaches to large language model-based "synthetic personas." The synthetic audiences article covers related media uses.
What it can and cannot do
| Can | Cannot |
|---|---|
| Smooth estimates for small subgroups | Add new information beyond the real sample |
| Help design follow-up research | Replace real respondents for decisions |
| Explore "what if" scenarios | Correct bias in the original sample |
| Speed early planning | Provide confident statistics for tiny groups |
If your real sample has 12 rare-disease specialists, generating 200 synthetic ones does not give you 212 opinions. It gives you 12 opinions, modeled.
Validation
Ask any provider:
- How do you validate synthetic results? Against held-out real responses?
- How close were predictions to real data in past studies, by question type?
- How do you handle subgroups with very few real respondents?
- How is uncertainty reported?
A good provider can show cases where synthetic results disagreed with real data and explain why.
Bias risk
Models reproduce and can amplify biases in training data. If the real sample overrepresents academic specialists, synthetic expansion will too. The AI audience model bias audit article explains how to check.
Appropriate uses
- Early hypothesis generation before a full study.
- Testing questionnaire design.
- Planning sample sizes for real research.
- Scenario exploration, clearly labeled.
Inappropriate uses
- Presenting synthetic findings as real HCP or patient views.
- Using synthetic data as evidence in promotional claims.
- Making major budget decisions on synthetic subgroups alone.
Common mistakes
- Reporting synthetic sample size as real sample size.
- Not disclosing that results are synthetic.
- Skipping validation against real data.
Practical takeaway
If a provider offers synthetic expansion, ask them to predict answers for a subgroup you already have real data on, then compare. Their accuracy on known data is the best guide to how much to trust them on unknown data.
Frequently asked questions
What is synthetic sample expansion?
Creating simulated survey respondents from a model trained on real responses, to increase the apparent size of small subgroups for analysis.
Is synthetic sample accurate?
It reproduces patterns in the training data. It can be useful for exploration, but it cannot discover views the real sample did not capture, and errors in the model carry through.
Should synthetic respondents be used for regulatory or medical decisions?
No. Use them, if at all, for early exploration and planning, clearly labeled, with real data for decisions.
Sources
External guidance and platform documentation change. Links were current at publication; check them again before relying on them for a decision.
Editorial note. Analysis and frameworks are the author's own and do not represent Acxiom or any current or former employer, client, or named platform. Examples labeled hypothetical or illustrative are not results from real campaigns. Nothing here is legal, regulatory, or medical advice.
New pharma programmatic breakdowns, occasionally
One email when I publish something worth reading. Benchmarks, measurement teardowns, and case studies with the caveats attached. No cadence promises, no reselling your address.
Unsubscribe any time. See the privacy policy.
Working through this decision on a real plan?
I work on health and pharma data, identity, and activation, after five years running HCP and DTC programmatic agency-side. Happy to talk through how this applies to your situation.