Audience Data Accuracy Validation: What Third-Party Truth Sets Tell You
How third party truth sets such as Truthset validate audience data accuracy, what accuracy scores mean, and where the method falls short for health data.
The short answer
Third-party truth sets validate audience data by comparing a vendor's attribute assignments, such as age, gender, or household income, against an independent reference source and scoring how often they agree. Truthset is one company that does this for consumer data. The scores are useful for comparing vendors on demographics, but they say much less about health condition segments, which are hard to verify without regulated data.
Every audience vendor says its data is accurate. Very few will show you how they know. Truth sets are one of the few ways to get an outside opinion, and they have become a regular part of data procurement for consumer brands. In pharma they are useful too, with some caveats that matter more here than in retail.
The reason this belongs in a measurement series: if the audience was wrong, the outcome study is measuring the wrong people. A weak Rx lift result can come from bad targeting as easily as from bad creative. The outcomes measurement platforms guide covers the downstream half of that problem.
How truth sets work
A truth set is a reference dataset believed to be accurate for certain attributes, usually built from verified or self-reported sources. Validation works like this:
- A data vendor submits its records with assigned attributes (for example, "female, 45 to 54, household income over $100,000").
- The validator matches those records to its reference data through a privacy-safe identity process.
- For each matched record, it compares the vendor's assignment with the reference.
- It reports accuracy, often as a score or index, by attribute and sometimes by record.
Truthset, founded in 2019, describes its approach as scoring the likelihood that a record and attribute are true, using a calibration panel and data from many providers, refreshed on a regular cycle. Its public materials describe tiered ratings that let buyers choose segments by accuracy level. Check Truthset's current documentation for exact methods; this summary is general.
What audience accuracy scores tell you
An accuracy score answers a narrow question: how often does this vendor's attribute assignment agree with the reference? That is valuable. It lets you compare vendors on the same attribute and spot segments that are close to random.
| Score tells you | Score does not tell you |
|---|---|
| Relative accuracy of a vendor on a given attribute | Whether the segment will perform in your campaign |
| Which segments are near random | Accuracy of attributes the truth set does not cover |
| Accuracy at a point in time | How fast the data decays after the test |
| Agreement with the reference source | Whether the reference itself is complete or biased |
A simple way to think about the stakes, using illustrative numbers: if you buy a 1,000,000-person segment and 50 percent of the assignments are correct, you are paying for 500,000 people outside the target. If the segment is 80 percent accurate but half the size, you reach 400,000 correct people instead of 500,000, with far less waste. Whether that trade is worth it depends on price and the cost of reaching the wrong person, which in pharma includes compliance exposure.
Limits of truth sets for health data
This is where pharma differs from consumer packaged goods. Most truth sets are strongest on demographics, because those can be verified from many sources. Health condition and treatment attributes are another matter.
- Verification is hard. Confirming that someone has a condition usually requires health data that is regulated under HIPAA or state laws. A general consumer truth set may not have it, and should not.
- Condition segments are usually modeled. Many health audiences are propensity models, not lists of diagnosed people. A truth set can validate the demographic inputs but not always the model's condition prediction.
- State laws narrow what is possible. Washington's My Health My Data Act and similar laws restrict collection and sharing of consumer health data, which limits both audience building and validation. See state consumer health data laws and pharma media.
- HCP data is a different problem. For HCP audiences, the reference is often the NPPES registry plus licensing and practice data. Accuracy questions are about specialty, practice location, and identity resolution, covered in HCP identity matching.
How to use validation in a data purchase
A numbered checklist for the next audience RFP:
- Ask each vendor whether its data has been independently validated, by whom, and on which attributes.
- Request the date of the most recent validation. Old scores describe old data. See data freshness in healthcare audiences.
- For health segments, ask separately how the condition flag was built and validated. Accept "modeled" as an answer, but ask what the model was trained on.
- Run your own check where you can: overlap a vendor segment with your CRM or a known list and see whether the attributes agree. Audience overlap analysis describes how.
- Compare scale and accuracy together. A table of segment size, accuracy, and cost per accurate record makes the trade visible.
- Write minimum accuracy expectations into the contract where it makes sense, using the health audience data contract checklist.
Connecting validation to outcome measurement
If you have both an accuracy score and an outcome study, read them together. A segment with low demographic accuracy and weak Rx lift probably has a targeting problem. A segment with high accuracy and weak lift points more toward creative, frequency, or channel. That is a cleaner diagnosis than blaming the agency or the publisher, and it is the kind of note that should appear in your monthly performance report.
Practical takeaway
Pick the single largest audience segment in your current DTC plan and ask the vendor, in writing, for the most recent independent accuracy validation on its core attributes and a plain description of how the health condition flag was built. If they cannot answer both, treat that segment as unvalidated and size your test budget accordingly.
Frequently asked questions
What is Truthset?
Truthset is a US company, founded in 2019, that scores the accuracy of consumer data at the record and attribute level. Data providers submit their data, and Truthset compares it against reference data to estimate how likely each assignment is to be true.
What does an audience accuracy score mean?
It is an estimate of how often an attribute assigned by a data vendor (for example, age range or gender) matches a reference source. A higher score means the vendor's assignments agree more often with the truth set, within the limits of that truth set.
Can truth sets validate health condition audiences?
Only partly. Health conditions are hard to verify against a reference without using regulated health data, so validation is usually stronger for demographic attributes than for diagnoses or treatment status. For health segments, ask the vendor how the condition flag itself was validated.
Should I only buy the highest-rated audience segments?
Not automatically. Higher accuracy usually means smaller scale, and the right tradeoff depends on what the campaign needs. For tightly regulated or narrow audiences, accuracy matters more; for broad awareness, a moderate score with more reach may be acceptable.
Sources
- Truthset, Data Ratings
- Washington State Attorney General, Protecting Washingtonians' Personal Health Data and Privacy
- FTC, Health Breach Notification Rule: The Basics for Business
External guidance and platform documentation change. Links were current at publication; check them again before relying on them for a decision.
Editorial note. Analysis and frameworks are the author's own and do not represent Acxiom or any current or former employer, client, or named platform. Examples labeled hypothetical or illustrative are not results from real campaigns. Nothing here is legal, regulatory, or medical advice.
New pharma programmatic breakdowns, occasionally
One email when I publish something worth reading. Benchmarks, measurement teardowns, and case studies with the caveats attached. No cadence promises, no reselling your address.
Unsubscribe any time. See the privacy policy.
Working through this decision on a real plan?
I work on health and pharma data, identity, and activation, after five years running HCP and DTC programmatic agency-side. Happy to talk through how this applies to your situation.