
Overview
This report presents a framework for evaluating and correcting representativeness failures in healthcare claims data used for AI and machine learning models in actuarial applications. It is accompanied by a model card generator and tool for comparing models.
Key Findings
Through a structured three‑stage framework and controlled synthetic scenarios, the report demonstrates that models can achieve strong aggregate performance while systematically failing subgroups affected by access barriers, documentation bias, or intersectional disadvantage. Based on the framework and results from the three scenarios in this report, it provides illustrates applications for actuaries working with machine learning models on claims data.
Podcast
Acknowledgements
The researchers’ deepest gratitude goes to those without whose efforts this project could not have come to fruition: the Project Oversight Group and others for their diligent work overseeing the research framework and deliverables, providing substantive feedback on methodology and findings, and reviewing and editing this report for accuracy and relevance.
Thanks to Sara Teppema, FSA, MAAA, FCA, who was an Advisor for the project team and a reviewer for the report and project deliverables.
Project Oversight Group members
Dorothy Andrews, ASA, MAAA, PhD
Andrew Dillworth, FSA, MAAA
Robert Gomez, FSA, MAAA, CERA
Matthew Kelsey, ASA
Eileen Luxton, FSA, FCIA
Marilyn McGaffin, ASA, MAAA, FLMI
Tri Pham, FSA, MAAA
Tony Pistilli, FSA, MAAA, CERA, FCA
Rose Qian, FSA, CERA
Becky Sheppard, FSA, MAAA
Justus Tulowiecki, ASA
At the Society of Actuaries Research Institute
Rob Montgomery ASA, MAAA, FLMI
Lisa Schilling, FSA, EA, FCA, MAAA
Barbara Scott, Senior Research Administrator
FAQ
Populations facing barriers to health care, such as rural residents, those with limited income, and those systematically underdiagnosed in clinical settings, appear in claims data at rates that do not reflect their true disease burden. When predictive models are trained on such data, models may learn and amplify these gaps, producing outputs that may be accurate on average while systematically failing specific subgroups.
Evaluate representativeness by comparing the healthcare claims data with a clearly defined target population. Assess whether the data reflects the target population's demographic, geographic, and clinical characteristics, identify any material representation gaps, and determine whether those gaps could affect model performance or accurately reflect the true disease burden in the population the model is intended to serve.