Diagnostic data and precision medicine commercialization

How to Judge the Accuracy of Lab and Diagnostic Data for Targeting

How to judge the accuracy and relevance of diagnostic data for HCP targeting: lab coverage, physician attribution, latency, de-identification, and validation.

Christian Guerrero Published 5 min read Part 3 of 10

The short answer

To judge the accuracy of diagnostic data for targeting, check five things: which labs are covered and which are missing, how each test is attributed to an ordering or treating NPI, how long it takes for a test to appear in the data, how de-identification and small-cell suppression work, and whether the vendor has validated its counts against any known reference. Then test a sample yourself against accounts you already understand.

Buyers ask some version of "how does this vendor ensure the accuracy and relevance of its diagnostic data?" in nearly every evaluation. The answer in the deck is usually a slide about proprietary methods and quality checks. That slide tells you nothing. You need answers you can verify, and a way to check them on your own list.

This piece is part of the diagnostic data and precision medicine launch series. It is written for any vendor in the category; none of it describes a specific company.

What "accurate" means for targeting

For HCP targeting, a lab dataset does not need to capture every test in the country. It needs to rank HCPs correctly enough that your tiers are right, and it needs to be biased in ways you understand. A dataset that sees 40 percent of testing evenly across settings can be more useful than one that sees 70 percent concentrated in one lab network. Those numbers are illustrative, but the point holds: evenness of coverage often matters more than total coverage.

Relevance is a separate test. The data has to cover your test type (a specific companion diagnostic, a multi-gene panel, or comprehensive profiling), in your tumor type or condition, recently enough to reflect current behavior.

Lab coverage: the first question

Ask for coverage broken down, not as a single headline figure:

  • Share of estimated national volume for your test type, and how the denominator was estimated
  • Reference labs vs. hospital and academic in-house labs
  • Specialty molecular labs, which may run most comprehensive panels
  • Regional gaps, if lab partners are concentrated geographically
  • Changes over time: labs that joined or left in the last 24 months

A lab leaving the data source looks exactly like HCPs who stopped testing. If a vendor cannot tell you when its lab mix changed, you cannot read a trend from its data.

Attribution: ordering vs. treating physician

A lab requisition names an ordering provider. That is often, but not always, the physician who will choose the therapy. Pathologists may order reflex testing. Fellows order under attendings. Some records carry only a facility or group number. Ask:

  1. Which NPI field is used, and what share of records have a valid individual NPI?
  2. How are group and facility identifiers handled: dropped, kept at account level, or distributed to individuals?
  3. Does the vendor infer a treating physician, and if so, from what linked data and with what confidence?
  4. How are NPIs validated against NPPES for status and specialty?

Inferred treating-physician links can be useful, but they are a model, not an observation. Treat them as you would any modeled audience, which is covered in when to use modeled HCP lists.

Latency and refresh

Latency is the time from a test being ordered to it appearing in the file you activate. It stacks: lab reporting to the vendor, vendor processing and de-identification, delivery to you, and then onboarding into a DSP or CRM. A two-week vendor lag can easily become six weeks in the media audience.

Ask for the median and the tail. If 10 percent of records arrive three months late, your most recent month always looks low and then gets restated. That makes recent testing "drops" a recurring false alarm. The general issue is covered in how old is too old for healthcare audience data.

De-identification and suppression

Commercial lab data should be de-identified under HIPAA, through either the Safe Harbor method or expert determination as described in HHS guidance. Ask which method applies, who performed any expert determination, when it was last renewed, and what it permits. Small counts are commonly suppressed or bucketed to reduce re-identification risk. That is appropriate, and you need to know the rule: if anything below a threshold is shown as zero or as a band, low testers and suppressed testers look the same.

The compliance side is covered more fully in patient identification with diagnostic data.

Validation against something you know

Ask the vendor what it has validated against. Good answers name a reference: lab partner totals, published testing rates in a defined population, or a reconciliation with claims procedure codes for the same tests. Weak answers describe internal consistency checks only. Then run your own check:

CheckHow to run itWhat a problem looks like
Known accountsPick 10 to 20 accounts your field team knows well; compare tiersKnown high testers flagged low, usually a coverage gap
Specialty mixTabulate testing NPIs by specialtyToo many pathologists or unknown specialties at the top
Trend stabilityCompare last three refreshesLarge swings that match lab additions, not behavior
Claims reconciliationCompare to procedure-code counts in claims if you have themRatios that vary wildly by region or setting
Recent-month restatementCompare the same month across two deliveriesLarge upward revisions, a sign of long-tail latency

The general approach to truth sets is in audience data accuracy validation.

Practical takeaway

Send every diagnostic data vendor on your shortlist the same written request: coverage by lab type for your specific test, the NPI attribution rule, median and 90th percentile latency, the suppression threshold, and a sample file for 20 accounts you choose. Score the answers side by side. A vendor that will not answer in writing has answered.

Frequently asked questions

How can you tell if diagnostic data is accurate?

Ask for coverage by lab type, the attribution method, and the typical latency for your test, then check a sample against things you know, such as accounts your field team knows well. Accuracy for targeting means the right HCPs are flagged with roughly the right relative volume, not that every test is captured.

What is the biggest accuracy risk in lab data?

Uneven coverage. If the data includes most reference lab volume but little hospital in-house lab volume, HCPs at large health systems will look like low testers. That bias carries into targeting and into measurement.

Does de-identification reduce accuracy?

It can. Expert determination and Safe Harbor methods remove or generalize identifiers, and small counts may be suppressed. That is the right trade for privacy, but you should know how suppression works so you read low counts correctly.

Sources

External guidance and platform documentation change. Links were current at publication; check them again before relying on them for a decision.

Editorial note. Analysis and frameworks are the author's own and do not represent Acxiom or any current or former employer, client, or named platform. Examples labeled hypothetical or illustrative are not results from real campaigns. Nothing here is legal, regulatory, or medical advice.

Working through this decision on a real plan?

I work on health and pharma data, identity, and activation, after five years running HCP and DTC programmatic agency-side. Happy to talk through how this applies to your situation.