Impact of class imbalance on chest x-ray classifiers: towards better evaluation practices for discrimination and calibration performance

Mosquera, Candelaria; Ferrer, Luciana; Milone, Diego; Luna, Daniel; Ferrante, Enzo

Computer Science > Computer Vision and Pattern Recognition

arXiv:2112.12843 (cs)

[Submitted on 23 Dec 2021 (v1), last revised 14 Mar 2022 (this version, v2)]

Title:Impact of class imbalance on chest x-ray classifiers: towards better evaluation practices for discrimination and calibration performance

Authors:Candelaria Mosquera, Luciana Ferrer, Diego Milone, Daniel Luna, Enzo Ferrante

View PDF

Abstract:This work aims to analyze standard evaluation practices adopted by the research community when assessing chest x-ray classifiers, particularly focusing on the impact of class imbalance in such appraisals. Our analysis considers a comprehensive definition of model performance, covering not only discriminative performance but also model calibration, a topic of research that has received increasing attention during the last years within the machine learning community. Firstly, we conducted a literature study to analyze common scientific practices and confirmed that: (1) even when dealing with highly imbalanced datasets, the community tends to use metrics that are dominated by the majority class; and (2) it is still uncommon to include calibration studies for chest x-ray classifiers, albeit its importance in the context of healthcare. Secondly, we perform a systematic experiment on two major chest x-ray datasets to explore the behavior of several performance metrics under different class ratios and show that widely adopted metrics can conceal the performance in the minority class. Finally, we recommend the inclusion of complementary metrics to better reflect the system's performance in such scenarios. Our study indicates that current evaluation practices adopted by the research community for chest x-ray computer-aided diagnosis systems may not reflect their performance in real clinical scenarios, and suggest alternatives to improve this situation.

Comments:	Conference on Health, Inference, and Learning (CHIL) 2022 - Invited non-archival presentation
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2112.12843 [cs.CV]
	(or arXiv:2112.12843v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2112.12843

Submission history

From: Candelaria Mosquera [view email]
[v1] Thu, 23 Dec 2021 20:57:47 UTC (3,941 KB)
[v2] Mon, 14 Mar 2022 12:29:28 UTC (3,942 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Impact of class imbalance on chest x-ray classifiers: towards better evaluation practices for discrimination and calibration performance

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Impact of class imbalance on chest x-ray classifiers: towards better evaluation practices for discrimination and calibration performance

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators