Published on: August 5, 2026

Evaluating Representativeness in Healthcare Claims Data: A Framework for AI Fairness in Actuarial Applications

Authors: Robert Lieberthal; Stacy Chen, MPH; Jawand Singh

soa-research-societal-purpose-logo.png

Overview

This report presents a framework for evaluating and correcting representativeness failures in healthcare claims data used for AI and machine learning models in actuarial applications. It is accompanied by a model card generator and tool for comparing models.

Key Findings

Through a structured three‑stage framework and controlled synthetic scenarios, the report demonstrates that models can achieve strong aggregate performance while systematically failing subgroups affected by access barriers, documentation bias, or intersectional disadvantage. Based on the framework and results from the three scenarios in this report, it provides illustrates applications for actuaries working with machine learning models on claims data.

Podcast

Acknowledgements

The researchers’ deepest gratitude goes to those without whose efforts this project could not have come to fruition: the Project Oversight Group and others for their diligent work overseeing the research framework and deliverables, providing substantive feedback on methodology and findings, and reviewing and editing this report for accuracy and relevance.

Thanks to Sara Teppema, FSA, MAAA, FCA, who was an Advisor for the project team and a reviewer for the report and project deliverables.

Project Oversight Group members

Dorothy Andrews, ASA, MAAA, PhD
Andrew Dillworth, FSA, MAAA
Robert Gomez, FSA, MAAA, CERA
Matthew Kelsey, ASA
Eileen Luxton, FSA, FCIA
Marilyn McGaffin, ASA, MAAA, FLMI
Tri Pham, FSA, MAAA
Tony Pistilli, FSA, MAAA, CERA, FCA
Rose Qian, FSA, CERA
Becky Sheppard, FSA, MAAA
Justus Tulowiecki, ASA

At the Society of Actuaries Research Institute

Rob Montgomery ASA, MAAA, FLMI
Lisa Schilling, FSA, EA, FCA, MAAA
Barbara Scott, Senior Research Administrator

FAQ

Populations facing barriers to health care, such as rural residents, those with limited income, and those systematically underdiagnosed in clinical settings, appear in claims data at rates that do not reflect their true disease burden. When predictive models are trained on such data, models may learn and amplify these gaps, producing outputs that may be accurate on average while systematically failing specific subgroups.

Evaluate representativeness by comparing the healthcare claims data with a clearly defined target population. Assess whether the data reflects the target population's demographic, geographic, and clinical characteristics, identify any material representation gaps, and determine whether those gaps could affect model performance or accurately reflect the true disease burden in the population the model is intended to serve.

Suggested Citation

Lieberthal, Robert D., Stacy Chen, and Jawand Singh. Evaluating Representativeness in Healthcare Claims Data. Society of Actuaries Research Institute, July 2026.   https://www.soa.org/resources/research-reports/2026/dei127-represent-hc-data/

Questions or Comments?

Give us your feedback! Take a short survey on this report. Take Survey 

If you have comments or questions, please send an email to Research@soa.org

Authors: Robert Lieberthal; Stacy Chen, MPH; Jawand Singh
Published on: August 5, 2026
Back to Top