How to Judge the Accuracy of Lab and Diagnostic Data for Targeting
How to judge the accuracy and relevance of diagnostic data for HCP targeting: lab coverage, physician attribution, latency, de-identification, and validation.
The short answer
To judge the accuracy of diagnostic data for targeting, check five things: which labs are covered and which are missing, how each test is attributed to an ordering or treating NPI, how long it takes for a test to appear in the data, how de-identification and small-cell suppression work, and whether the vendor has validated its counts against any known reference. Then test a sample yourself against accounts you already understand.
Buyers ask some version of "how does this vendor ensure the accuracy and relevance of its diagnostic data?" in nearly every evaluation. The answer in the deck is usually a slide about proprietary methods and quality checks. That slide tells you nothing. You need answers you can verify, and a way to check them on your own list.
This piece is part of the diagnostic data and precision medicine launch series. It is written for any vendor in the category; none of it describes a specific company.
What "accurate" means for targeting
For HCP targeting, a lab dataset does not need to capture every test in the country. It needs to rank HCPs correctly enough that your tiers are right, and it needs to be biased in ways you understand. A dataset that sees 40 percent of testing evenly across settings can be more useful than one that sees 70 percent concentrated in one lab network. Those numbers are illustrative, but the point holds: evenness of coverage often matters more than total coverage.
Relevance is a separate test. The data has to cover your test type (a specific companion diagnostic, a multi-gene panel, or comprehensive profiling), in your tumor type or condition, recently enough to reflect current behavior.
Lab coverage: the first question
Ask for coverage broken down, not as a single headline figure:
- Share of estimated national volume for your test type, and how the denominator was estimated
- Reference labs vs. hospital and academic in-house labs
- Specialty molecular labs, which may run most comprehensive panels
- Regional gaps, if lab partners are concentrated geographically
- Changes over time: labs that joined or left in the last 24 months
A lab leaving the data source looks exactly like HCPs who stopped testing. If a vendor cannot tell you when its lab mix changed, you cannot read a trend from its data.
Attribution: ordering vs. treating physician
A lab requisition names an ordering provider. That is often, but not always, the physician who will choose the therapy. Pathologists may order reflex testing. Fellows order under attendings. Some records carry only a facility or group number. Ask:
- Which NPI field is used, and what share of records have a valid individual NPI?
- How are group and facility identifiers handled: dropped, kept at account level, or distributed to individuals?
- Does the vendor infer a treating physician, and if so, from what linked data and with what confidence?
- How are NPIs validated against NPPES for status and specialty?
Inferred treating-physician links can be useful, but they are a model, not an observation. Treat them as you would any modeled audience, which is covered in when to use modeled HCP lists.
Latency and refresh
Latency is the time from a test being ordered to it appearing in the file you activate. It stacks: lab reporting to the vendor, vendor processing and de-identification, delivery to you, and then onboarding into a DSP or CRM. A two-week vendor lag can easily become six weeks in the media audience.
Ask for the median and the tail. If 10 percent of records arrive three months late, your most recent month always looks low and then gets restated. That makes recent testing "drops" a recurring false alarm. The general issue is covered in how old is too old for healthcare audience data.
De-identification and suppression
Commercial lab data should be de-identified under HIPAA, through either the Safe Harbor method or expert determination as described in HHS guidance. Ask which method applies, who performed any expert determination, when it was last renewed, and what it permits. Small counts are commonly suppressed or bucketed to reduce re-identification risk. That is appropriate, and you need to know the rule: if anything below a threshold is shown as zero or as a band, low testers and suppressed testers look the same.
The compliance side is covered more fully in patient identification with diagnostic data.
Validation against something you know
Ask the vendor what it has validated against. Good answers name a reference: lab partner totals, published testing rates in a defined population, or a reconciliation with claims procedure codes for the same tests. Weak answers describe internal consistency checks only. Then run your own check:
| Check | How to run it | What a problem looks like |
|---|---|---|
| Known accounts | Pick 10 to 20 accounts your field team knows well; compare tiers | Known high testers flagged low, usually a coverage gap |
| Specialty mix | Tabulate testing NPIs by specialty | Too many pathologists or unknown specialties at the top |
| Trend stability | Compare last three refreshes | Large swings that match lab additions, not behavior |
| Claims reconciliation | Compare to procedure-code counts in claims if you have them | Ratios that vary wildly by region or setting |
| Recent-month restatement | Compare the same month across two deliveries | Large upward revisions, a sign of long-tail latency |
The general approach to truth sets is in audience data accuracy validation.
Practical takeaway
Send every diagnostic data vendor on your shortlist the same written request: coverage by lab type for your specific test, the NPI attribution rule, median and 90th percentile latency, the suppression threshold, and a sample file for 20 accounts you choose. Score the answers side by side. A vendor that will not answer in writing has answered.
Frequently asked questions
How can you tell if diagnostic data is accurate?
Ask for coverage by lab type, the attribution method, and the typical latency for your test, then check a sample against things you know, such as accounts your field team knows well. Accuracy for targeting means the right HCPs are flagged with roughly the right relative volume, not that every test is captured.
What is the biggest accuracy risk in lab data?
Uneven coverage. If the data includes most reference lab volume but little hospital in-house lab volume, HCPs at large health systems will look like low testers. That bias carries into targeting and into measurement.
Does de-identification reduce accuracy?
It can. Expert determination and Safe Harbor methods remove or generalize identifiers, and small counts may be suppressed. That is the right trade for privacy, but you should know how suppression works so you read low counts correctly.
Sources
- U.S. Department of Health and Human Services, Guidance Regarding Methods for De-identification of PHI
- Centers for Medicare and Medicaid Services, Clinical Laboratory Improvement Amendments (CLIA)
- CMS, NPPES NPI Registry
External guidance and platform documentation change. Links were current at publication; check them again before relying on them for a decision.
Editorial note. Analysis and frameworks are the author's own and do not represent Acxiom or any current or former employer, client, or named platform. Examples labeled hypothetical or illustrative are not results from real campaigns. Nothing here is legal, regulatory, or medical advice.
New pharma programmatic breakdowns, occasionally
One email when I publish something worth reading. Benchmarks, measurement teardowns, and case studies with the caveats attached. No cadence promises, no reselling your address.
Unsubscribe any time. See the privacy policy.
Working through this decision on a real plan?
I work on health and pharma data, identity, and activation, after five years running HCP and DTC programmatic agency-side. Happy to talk through how this applies to your situation.