A survey of bias in Machine Learning through the prism of Statistical Parity for the Adult Data Set

Besse, Philippe; del Barrio, Eustasio; Gordaliza, Paula; Loubes, Jean-Michel; Risser, Laurent

Statistics > Machine Learning

arXiv:2003.14263 (stat)

[Submitted on 31 Mar 2020 (v1), last revised 6 Apr 2020 (this version, v2)]

Title:A survey of bias in Machine Learning through the prism of Statistical Parity for the Adult Data Set

Authors:Philippe Besse, Eustasio del Barrio, Paula Gordaliza, Jean-Michel Loubes, Laurent Risser

View PDF

Abstract:Applications based on Machine Learning models have now become an indispensable part of the everyday life and the professional world. A critical question then recently arised among the population: Do algorithmic decisions convey any type of discrimination against specific groups of population or minorities? In this paper, we show the importance of understanding how a bias can be introduced into automatic decisions. We first present a mathematical framework for the fair learning problem, specifically in the binary classification setting. We then propose to quantify the presence of bias by using the standard Disparate Impact index on the real and well-known Adult income data set. Finally, we check the performance of different approaches aiming to reduce the bias in binary classification outcomes. Importantly, we show that some intuitive methods are ineffective. This sheds light on the fact trying to make fair machine learning models may be a particularly challenging task, in particular when the training observations contain a bias.

Subjects:	Machine Learning (stat.ML); Computers and Society (cs.CY); Machine Learning (cs.LG)
Cite as:	arXiv:2003.14263 [stat.ML]
	(or arXiv:2003.14263v2 [stat.ML] for this version)
	https://doi.org/10.48550/arXiv.2003.14263

Submission history

From: Paula Gordaliza [view email]
[v1] Tue, 31 Mar 2020 14:48:36 UTC (354 KB)
[v2] Mon, 6 Apr 2020 11:16:10 UTC (354 KB)

Statistics > Machine Learning

Title:A survey of bias in Machine Learning through the prism of Statistical Parity for the Adult Data Set

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Statistics > Machine Learning

Title:A survey of bias in Machine Learning through the prism of Statistical Parity for the Adult Data Set

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators