Synthetic Control and Matched Market Tests for Pharma Media
How geo tests, matched market tests, and synthetic control work for pharma media, when to use them instead of holdouts, and a hypothetical worked example.
The short answer
Matched market tests and synthetic control are geographic experiments: you run or change media in some markets and compare their outcomes with markets that did not get the change. A matched market test pairs markets that looked alike before the test. Synthetic control builds a weighted blend of many untreated markets that tracks the test markets closely in the pre-period, then uses that blend as the counterfactual. Pharma teams use these designs for TV, CTV, and other media that cannot be withheld from individual prescribers or patients.
A randomized holdout is the cleanest test, but it needs a list you can suppress. National TV does not have one. Neither does radio, most out-of-home, or a CTV buy targeted by household graph. For those channels the unit of the experiment becomes the market, usually a DMA, and the method becomes geographic.
The tradeoffs between prescriber-level and geographic designs are covered in HCP-level vs. geographic pharma measurement designs. This article explains the geographic methods in plain terms and walks through a hypothetical example.
What is a geo test?
A geo test assigns markets, rather than people, to test and control. Test markets get the media change (more spend, a new channel, or the campaign itself). Control markets do not. After the flight, you compare the change in outcomes, usually prescriptions from claims or syndicated data aggregated to the market.
Three common versions:
| Design | How control is built | Strength | Weakness |
|---|---|---|---|
| Randomized geo test | Markets randomly assigned to test or control | Unbiased in expectation | Needs many markets; random draws can be unbalanced |
| Matched market test | Each test market paired with a similar control market | Simple, easy to explain | Depends on the pairing; one odd market distorts the read |
| Synthetic control | Weighted blend of many donor markets fitted to the test markets' pre-period | Closer pre-period fit; uses more data | Harder to explain; sensitive to donor pool choices |
Matched market tests explained
You pick test markets, then find control markets that match them on the things that drive prescriptions: historical Rx trend, market size, payer mix, field force coverage, specialist density, and competitor presence. Pairs might look like one Midwest DMA matched with another of similar size and trend.
Matched markets are easy to explain in a QBR, which is a real advantage. The weakness is that matching on a few variables leaves room for differences you did not match on. A regional health system changing its formulary, a local news story, or a field territory realignment can hit one market in a pair and not the other.
Synthetic control explained simply
Synthetic control asks a different question. Instead of "which single market looks most like my test market," it asks "what weighted combination of all my untreated markets would have tracked my test markets most closely before the test?"
The model might find that 30 percent of DMA A, 25 percent of DMA B, 20 percent of DMA C, and smaller slices of several others reproduce the test markets' weekly prescription line during the pre-period almost exactly. That blend is the synthetic control. After launch, the blend continues on its path, and the gap between the real test markets and the blend is the estimated effect.
Because the weights come from data rather than a planner's judgement, the pre-period fit is usually tighter than a hand-picked match. Open-source tools such as Meta's GeoLift implement versions of this approach, and some MMM tools, including Google's Meridian, offer ways to bring geo experiment results into the model. Check the current documentation for what each supports.
A hypothetical worked example
All figures are hypothetical. A DTC brand wants to know whether a CTV campaign drives new patient starts. It picks 10 test DMAs and uses 40 other DMAs as the donor pool. CTV runs in the test DMAs for 8 weeks; the donor DMAs get no CTV. All other media stays constant nationally.
- Pre-period fit. Over the 26 weeks before launch, the 10 test DMAs average 2,000 NBRx per week. The synthetic control built from donor DMAs also averages 2,000 per week, with weekly differences small relative to the total.
- Post-period result. During the 8-week flight, test DMAs average 2,300 NBRx per week. The synthetic control averages 2,150 per week.
- Weekly effect. 2,300 minus 2,150 equals 150 incremental NBRx per week.
- Total effect. 150 times 8 weeks equals 1,200 incremental NBRx.
- Percent lift. 150 divided by 2,150 equals about 7.0 percent.
- Cost per incremental NBRx. At a hypothetical $1,200,000 CTV spend in test markets, $1,200,000 divided by 1,200 equals $1,000.
The next step is checking whether 150 per week is bigger than the normal noise. A common approach is a placebo test: pretend each donor market was treated, rerun the model, and see how often a gap that large appears by chance. If many placebo markets show gaps of 150 or more, the result is not convincing.
When to use geo tests in pharma
- Use them for linear TV, broad CTV, audio, out-of-home, and national DTC campaigns, or to test a new channel before rolling it out. For CTV specifics, see how to test incrementality in pharma CTV.
- Use them to calibrate a marketing mix model, since geo variation is the kind of signal MMM needs.
- Avoid them for HCP programmatic when you have an NPI list. An individual holdout test has far more units and therefore more power.
- Be careful with rare disease brands. Weekly prescriptions per market may be too few to read.
What usually goes wrong with geo tests
- Spillover. Digital media leaks across DMA borders, and patients travel to specialists in other markets. Exclude border-heavy markets or measure by where the patient lives.
- Field force changes. A territory realignment in the middle of the test can swamp the media effect. Coordinate with sales operations before launch.
- Too few markets. One local event in a small test set can drive the whole result.
- National media contamination. If the brand also runs national TV, the test is measuring CTV on top of TV, not CTV alone. Say so in the readout.
- No power analysis. Run the design on historical data first to see how large an effect you could actually detect.
Geo results should be reported the same way as any other study, with the design and the uncertainty next to the headline number, as set out in the outcomes measurement platforms guide.
Practical takeaway
Before running a geo test, take the last 52 weeks of NBRx by DMA and run a dry test: pick your candidate test markets, build the matched or synthetic control on the first 26 weeks, and check how far the control drifts from the test markets over the next 8 weeks with no media change at all. That drift is your noise floor, and if it is larger than the lift you hope to see, change the design before you spend.
Frequently asked questions
What is a matched market test?
A matched market test runs media in some geographic markets and withholds or changes it in similar markets, then compares outcomes. The markets are paired or grouped so they behaved alike before the test.
What is synthetic control in marketing?
Synthetic control builds a weighted blend of untreated markets that closely tracks the test markets during the pre-period. After launch, that blend serves as the estimate of what the test markets would have done without the media, and the gap is the estimated effect.
When should pharma use a geo test instead of a holdout?
Use a geo test when you cannot suppress media at the individual level, such as linear TV, broad CTV, radio, out-of-home, or national campaigns. Use an individual holdout when you have an addressable list, such as an NPI target list, because it gives more statistical power.
How many markets do I need for a geo test?
There is no single number, but more is better. With only a handful of test markets, one local event can swamp the result. Many teams use a set of test DMAs and a larger donor pool, and run a power analysis on historical data before launch.
Sources
- Google for Developers, Meridian
- Meta Open Source, GeoLift (GitHub)
- Media Rating Council, Standards and Guidelines
External guidance and platform documentation change. Links were current at publication; check them again before relying on them for a decision.
Editorial note. Analysis and frameworks are the author's own and do not represent Acxiom or any current or former employer, client, or named platform. Examples labeled hypothetical or illustrative are not results from real campaigns. Nothing here is legal, regulatory, or medical advice.
New pharma programmatic breakdowns, occasionally
One email when I publish something worth reading. Benchmarks, measurement teardowns, and case studies with the caveats attached. No cadence promises, no reselling your address.
Unsubscribe any time. See the privacy policy.
Working through this decision on a real plan?
I work on health and pharma data, identity, and activation, after five years running HCP and DTC programmatic agency-side. Happy to talk through how this applies to your situation.