Provable Detection of Propagating Sampling Bias in Prediction Models

Ravishankar, Pavan; Mo, Qingyu; McFowland III, Edward; Neill, Daniel B.

Computer Science > Machine Learning

arXiv:2302.06752 (cs)

[Submitted on 13 Feb 2023]

Title:Provable Detection of Propagating Sampling Bias in Prediction Models

Authors:Pavan Ravishankar, Qingyu Mo, Edward McFowland III, Daniel B. Neill

View PDF

Abstract:With an increased focus on incorporating fairness in machine learning models, it becomes imperative not only to assess and mitigate bias at each stage of the machine learning pipeline but also to understand the downstream impacts of bias across stages. Here we consider a general, but realistic, scenario in which a predictive model is learned from (potentially biased) training data, and model predictions are assessed post-hoc for fairness by some auditing method. We provide a theoretical analysis of how a specific form of data bias, differential sampling bias, propagates from the data stage to the prediction stage. Unlike prior work, we evaluate the downstream impacts of data biases quantitatively rather than qualitatively and prove theoretical guarantees for detection. Under reasonable assumptions, we quantify how the amount of bias in the model predictions varies as a function of the amount of differential sampling bias in the data, and at what point this bias becomes provably detectable by the auditor. Through experiments on two criminal justice datasets -- the well-known COMPAS dataset and historical data from NYPD's stop and frisk policy -- we demonstrate that the theoretical results hold in practice even when our assumptions are relaxed.

Comments:	AAAI 2023 (13 pages, 7 figures)
Subjects:	Machine Learning (cs.LG); Computers and Society (cs.CY)
Cite as:	arXiv:2302.06752 [cs.LG]
	(or arXiv:2302.06752v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2302.06752

Submission history

From: Pavan Ravishankar [view email]
[v1] Mon, 13 Feb 2023 23:39:35 UTC (8,146 KB)

Computer Science > Machine Learning

Title:Provable Detection of Propagating Sampling Bias in Prediction Models

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Provable Detection of Propagating Sampling Bias in Prediction Models

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators