No imputation without representation

Lenz, Oliver Urs; Peralta, Daniel; Cornelis, Chris

Computer Science > Machine Learning

arXiv:2206.14254 (cs)

[Submitted on 28 Jun 2022 (v1), last revised 25 Oct 2022 (this version, v3)]

Title:No imputation without representation

Authors:Oliver Urs Lenz, Daniel Peralta, Chris Cornelis

View PDF

Abstract:Imputation allows datasets to be used with algorithms that cannot handle missing values by themselves. However, missing values may in principle contribute useful information that is lost through imputation. The missing-indicator approach can be used to preserve this information. There are several theoretical considerations why missing-indicators may or may not be beneficial, but there has not been any large-scale practical experiment on real-life datasets to test this question for machine learning predictions. We perform this experiment for three imputation strategies and a range of different classification algorithms, on the basis of twenty real-life datasets. We find that missing-indicators generally increase classification performance, and that nearest neighbour and iterative imputation do not lead to better performance than simple mean/mode imputation. Therefore, we recommend the use of missing-indicators with mean/mode imputation as a safe default, with the caveat that for decision trees, pruning is necessary to prevent overfitting.

Subjects:	Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as:	arXiv:2206.14254 [cs.LG]
	(or arXiv:2206.14254v3 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2206.14254

Submission history

From: Oliver Urs Lenz [view email]
[v1] Tue, 28 Jun 2022 19:12:09 UTC (46 KB)
[v2] Sun, 16 Oct 2022 11:54:33 UTC (58 KB)
[v3] Tue, 25 Oct 2022 18:02:11 UTC (46 KB)

Computer Science > Machine Learning

Title:No imputation without representation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:No imputation without representation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators