Rx outcomes measurement platforms and methods

What Is a Holdout Test? A Plain-Language Guide for Marketers

What is a holdout test? A plain explanation of how marketers withhold ads from a random group, size the test, read the result, and avoid the common mistakes.

Christian Guerrero Published 6 min read Part 6 of 10

The short answer

A holdout test is an experiment where you randomly pick part of your audience and do not show them your ads, then compare their outcomes with the group that could see ads. Because the split is random, the difference between the two groups estimates the true effect of the advertising. It is the simplest reliable way to answer "would this have happened anyway?"

Most marketing reports tell you what happened among people who saw your ads. They do not tell you what would have happened if those people had not seen them. A holdout test is how you find out. It is an old idea, borrowed from clinical trials and direct mail, and it still beats most of the fancier methods when you can run it.

This is the beginner version. For pharma-specific design, contamination rules, and power calculations, read how to design a holdout test for pharma media.

What is a holdout test and why does it work?

You start with a defined audience: a CRM file, a target list of physicians, a set of households. Before the campaign starts, you randomly assign each member to one of two groups.

  • Test group: eligible to see the ads.
  • Holdout group (control): blocked from seeing the ads.

Random assignment is the whole trick. If the split is random and the groups are large enough, they will be alike on average in every way that matters: age, behavior, prescribing habits, loyalty, even things you never measured. So any difference in outcomes after the campaign is the campaign's doing.

Compare that with an exposed vs. unexposed comparison. People who happen to see ads are different from people who do not. They are online more, they visited your site, they are in a retargeting pool. A naive comparison credits the ad with differences that were already there. The difference between attributed and incremental results comes from exactly this.

How to set up a holdout test

  1. Define the audience. A fixed list works best, such as an NPI target list or a set of hashed emails.
  2. Pick the outcome. One primary metric: a sale, a sign-up, a new prescription. Decide this before launch.
  3. Randomize. Use a random number, not alphabetical order or region. Many teams stratify first (for example, randomize within each prescriber decile) so both groups have the same mix.
  4. Suppress the holdout. Upload the holdout as an exclusion list in every platform running the tactic. Check that it is applied.
  5. Run the campaign. Do not change the holdout midway.
  6. Measure both groups the same way. Same data source, same window, same definitions.
  7. Compare. Difference in outcome rate, with a confidence interval.

Sizing basics: how big should the holdout be?

Bigger holdouts give more precise answers but cost more reach. Common practice is 10 to 20 percent. What actually decides it:

FactorEffect on holdout size
Small total audienceNeeds a larger share held out, or a longer test
Low baseline outcome rateNeeds more people to see a difference
Small expected effectNeeds more people; small effects hide in noise
High cost of lost reach (launch, competitive pressure)Pushes toward a smaller holdout and longer test

A hypothetical example. A target list has 20,000 physicians. You hold out 15 percent: 3,000 in holdout, 17,000 in test. Over the test period, 4.4 percent of the test group write a new prescription for the brand, versus 4.0 percent of the holdout. The difference is 0.4 percentage points, which is a 10 percent relative lift (0.4 divided by 4.0). Across the 17,000 test-group physicians, that is about 68 incremental prescribers (0.004 times 17,000). Whether 0.4 points is distinguishable from noise depends on the confidence interval, which your analyst or measurement partner should calculate. A significant result can still be commercially small, which is covered in statistical significance vs. commercial importance.

Note that the comparison is between everyone assigned to test and everyone assigned to holdout, not only those who actually saw an ad. That keeps the randomization intact. Analysts call this intent-to-treat.

Common holdout test mistakes

  • Leaky holdouts. The holdout is suppressed in one DSP but not in another, or the same people see your ads on a publisher direct buy. Check every channel running the tactic.
  • Comparing only the exposed. Dropping test-group members who never saw an ad reintroduces the bias randomization removed.
  • Stopping early. Peeking weekly and stopping when the number looks good produces false winners.
  • Changing the outcome after launch. If the primary metric shifts because the first one looked flat, the test is no longer a test.
  • Too small to read. A holdout of a few hundred people with a rare outcome will almost always come back inconclusive.
  • Treating one result as permanent. Media, creative, and competitors change. Retest the big tactics periodically.

When a holdout test is not the right tool

Holdouts need an audience you can address and suppress at the individual level. For linear TV, out-of-home, radio, or broad CTV, you usually cannot block specific people, so teams use geographic designs instead. That approach is covered in synthetic control and matched market tests. For long-run cross-channel budget questions, marketing mix modeling is the more useful tool, ideally calibrated with holdout results.

In pharma, a holdout on an HCP target list is one of the cleanest ways to attribute prescription lift to programmatic HCP media, and it is the check I would run on any always-on attribution number that drives a large budget decision.

Practical takeaway

For your next campaign with a fixed audience file, randomly assign 15 percent of it to a holdout before launch, upload that holdout as an exclusion list in every platform, and write down the single outcome you will compare. That setup takes an afternoon and gives you a result you can defend.

Frequently asked questions

What is a holdout test in marketing?

A holdout test randomly withholds advertising from part of the target audience and compares outcomes between the people who could see ads and the people who could not. The difference estimates how much the advertising caused.

How big should a holdout group be?

Common holdout sizes range from about 10 to 20 percent of the audience, but the right size depends on audience volume, the baseline outcome rate, and the size of effect you need to detect. Small audiences may need a larger share held out or a longer test.

Does a holdout test waste budget?

It reduces reach to the held-out group for the test period, which has an opportunity cost. In exchange you get a causal read on whether the spend works, which often saves far more than it costs if the tactic turns out not to be incremental.

What is the difference between a holdout test and an A/B test?

An A/B test compares two versions of an ad or experience. A holdout test compares some advertising with none. Both rely on random assignment.

Sources

External guidance and platform documentation change. Links were current at publication; check them again before relying on them for a decision.

Editorial note. Analysis and frameworks are the author's own and do not represent Acxiom or any current or former employer, client, or named platform. Examples labeled hypothetical or illustrative are not results from real campaigns. Nothing here is legal, regulatory, or medical advice.

Working through this decision on a real plan?

I work on health and pharma data, identity, and activation, after five years running HCP and DTC programmatic agency-side. Happy to talk through how this applies to your situation.