AI applications and emerging healthcare media developments

How to Audit Bias in an AI-Assisted Healthcare Audience Model

A subgroup evaluation plan for AI-assisted healthcare audience models, tied to intended use, representativeness, missing data, and drift.

Christian Guerrero Published 3 min read Part 8 of 10

The short answer

AI models increasingly decide which people see healthcare ads. A model trained on past data can reproduce patterns in that data, including gaps in who was diagnosed, treated, or reached. The result can be audiences that underrepresent some groups who could benefit from information about a condition or treatment. A bias audit checks whether the model works fairly across relevant groups for its intended use.

Define the intended use first

A model built to find people likely to have a condition should be evaluated on that task. Start by writing:

  • What the model predicts.
  • What decision uses the prediction.
  • Which groups matter for fairness in this context.
  • What harm could result from unequal performance.

Choose subgroups carefully

Relevant subgroups might include age bands, sex, geographic region, urban versus rural, and other characteristics relevant to the condition. Use subgroups for which you have lawful and appropriate data. Do not collect new sensitive data just to run the audit without privacy review.

Metrics to compare across subgroups

Metric Question
Selection rate What share of each group does the model select?
Precision Of those selected, how many truly fit the target?
Recall Of those who truly fit, how many were selected?
Calibration Do predicted scores mean the same thing across groups?

Large differences in recall are often the most important in healthcare marketing, because they mean some groups who could benefit are being missed.

Check representativeness of training data

  • Which groups were well represented in the data used to build the model?
  • Were some groups underdiagnosed or undertreated historically, making them look less likely to fit?
  • Is missing data concentrated in certain groups?

If training data reflects historical gaps, the model may reinforce them.

Watch for drift

Model performance can change as populations and data sources change. Repeat the audit periodically, especially after data refreshes or model updates.

What to do with findings

  • Adjust the model or its thresholds to reduce gaps.
  • Supplement with other approaches, such as contextual targeting, for underreached groups.
  • Document known limitations for everyone using the audience.

The NIST AI Risk Management Framework includes guidance on managing bias as part of broader AI risk.

Ask vendors

If the model comes from a vendor, ask for their bias evaluation, the subgroups tested, and the results. A vendor that has not evaluated subgroup performance is not ready for healthcare use. See synthetic audiences and data freshness.

Practical takeaway

Before using any AI-built audience at scale, compare selection rate and recall across at least three relevant subgroups. Document the results and revisit them after every model update.

Sources

External guidance and platform documentation change. Links were current at publication; check them again before relying on them for a decision.

Editorial note. Analysis and frameworks are the author's own and do not represent Acxiom or any current or former employer, client, or named platform. Examples labeled hypothetical or illustrative are not results from real campaigns. Nothing here is legal, regulatory, or medical advice.

Working through this decision on a real plan?

I work on health and pharma data, identity, and activation, after five years running HCP and DTC programmatic agency-side. Happy to talk through how this applies to your situation.