Reducing Sentiment Bias in Language Models via Counterfactual Evaluation

Huang, Po-Sen; Zhang, Huan; Jiang, Ray; Stanforth, Robert; Welbl, Johannes; Rae, Jack; Maini, Vishal; Yogatama, Dani; Kohli, Pushmeet

Computer Science > Computation and Language

arXiv:1911.03064v2 (cs)

[Submitted on 8 Nov 2019 (v1), revised 30 Apr 2020 (this version, v2), latest version 8 Oct 2020 (v3)]

Title:Reducing Sentiment Bias in Language Models via Counterfactual Evaluation

Authors:Po-Sen Huang, Huan Zhang, Ray Jiang, Robert Stanforth, Johannes Welbl, Jack Rae, Vishal Maini, Dani Yogatama, Pushmeet Kohli

View PDF

Abstract:Recent advances in language model architectures and the availability of large text corpora have driven progress on automatic text generation. While this results in models that are capable of generating coherent texts, it also prompts models to internalize social biases present in the training corpus. This paper aims to quantify and reduce a particular type of bias exhibited by language models: bias with respect to sentiment. Given a conditioning context (e.g. a writing prompt) and a language model, we analyze if (and how) the sentiment of the generated text is affected by changes in values of sensitive attributes (e.g. country names, occupations, genders) in the conditioning context using a form of counterfactual evaluation. We quantify bias by adopting individual and group fairness metrics from the fair machine learning literature, and demonstrate that large-scale models trained on two different corpora (news articles, and Wikipedia) exhibit considerable sentiment bias. We then propose the use of a sentiment prediction-derived regularization on the language model's latent representations. The regularization improves fairness metrics by 14--16% while retaining comparable levels of perplexity and semantic similarity.

Subjects:	Computation and Language (cs.CL); Computers and Society (cs.CY); Machine Learning (cs.LG)
Cite as:	arXiv:1911.03064 [cs.CL]
	(or arXiv:1911.03064v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.1911.03064

Submission history

From: Po-Sen Huang [view email]
[v1] Fri, 8 Nov 2019 05:56:01 UTC (146 KB)
[v2] Thu, 30 Apr 2020 17:51:20 UTC (644 KB)
[v3] Thu, 8 Oct 2020 17:58:35 UTC (690 KB)

Computer Science > Computation and Language

Title:Reducing Sentiment Bias in Language Models via Counterfactual Evaluation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Reducing Sentiment Bias in Language Models via Counterfactual Evaluation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators