An AI-Guided Data Centric Strategy to Detect and Mitigate Biases in Healthcare Datasets
Authors:
Faris F. Gulamali,
Ashwin S. Sawant,
Lora Liharska,
Carol R. Horowitz,
Lili Chan,
Patricia H. Kovatch,
Ira Hofer,
Karandeep Singh,
Lynne D. Richardson,
Emmanuel Mensah,
Alexander W Charney,
David L. Reich,
Jianying Hu,
Girish N. Nadkarni
Abstract:
The adoption of diagnosis and prognostic algorithms in healthcare has led to concerns about the perpetuation of bias against disadvantaged groups of individuals. Deep learning methods to detect and mitigate bias have revolved around modifying models, optimization strategies, and threshold calibration with varying levels of success. Here, we generate a data-centric, model-agnostic, task-agnostic ap…
▽ More
The adoption of diagnosis and prognostic algorithms in healthcare has led to concerns about the perpetuation of bias against disadvantaged groups of individuals. Deep learning methods to detect and mitigate bias have revolved around modifying models, optimization strategies, and threshold calibration with varying levels of success. Here, we generate a data-centric, model-agnostic, task-agnostic approach to evaluate dataset bias by investigating the relationship between how easily different groups are learned at small sample sizes (AEquity). We then apply a systematic analysis of AEq values across subpopulations to identify and mitigate manifestations of racial bias in two known cases in healthcare - Chest X-rays diagnosis with deep convolutional neural networks and healthcare utilization prediction with multivariate logistic regression. AEq is a novel and broadly applicable metric that can be applied to advance equity by diagnosing and remediating bias in healthcare datasets.
△ Less
Submitted 6 November, 2023;
originally announced November 2023.
Predicting Adverse Neonatal Outcomes for Preterm Neonates with Multi-Task Learning
Authors:
**gyang Lin,
Junyu Chen,
Hanjia Lyu,
Igor Khodak,
Divya Chhabra,
Colby L Day Richardson,
Irina Prelipcean,
Andrew M Dylag,
Jiebo Luo
Abstract:
Diagnosis of adverse neonatal outcomes is crucial for preterm survival since it enables doctors to provide timely treatment. Machine learning (ML) algorithms have been demonstrated to be effective in predicting adverse neonatal outcomes. However, most previous ML-based methods have only focused on predicting a single outcome, ignoring the potential correlations between different outcomes, and pote…
▽ More
Diagnosis of adverse neonatal outcomes is crucial for preterm survival since it enables doctors to provide timely treatment. Machine learning (ML) algorithms have been demonstrated to be effective in predicting adverse neonatal outcomes. However, most previous ML-based methods have only focused on predicting a single outcome, ignoring the potential correlations between different outcomes, and potentially leading to suboptimal results and overfitting issues. In this work, we first analyze the correlations between three adverse neonatal outcomes and then formulate the diagnosis of multiple neonatal outcomes as a multi-task learning (MTL) problem. We then propose an MTL framework to jointly predict multiple adverse neonatal outcomes. In particular, the MTL framework contains shared hidden layers and multiple task-specific branches. Extensive experiments have been conducted using Electronic Health Records (EHRs) from 121 preterm neonates. Empirical results demonstrate the effectiveness of the MTL framework. Furthermore, the feature importance is analyzed for each neonatal outcome, providing insights into model interpretability.
△ Less
Submitted 27 March, 2023;
originally announced March 2023.