Mitigating False-Negative Contexts in Multi-document QuestionAnswering with Retrieval Marginalization

Ni, Ansong; Gardner, Matt; Dasigi, Pradeep

Computer Science > Computation and Language

arXiv:2103.12235v1 (cs)

[Submitted on 22 Mar 2021 (this version), latest version 8 Sep 2021 (v2)]

Title:Mitigating False-Negative Contexts in Multi-document QuestionAnswering with Retrieval Marginalization

Authors:Ansong Ni, Matt Gardner, Pradeep Dasigi

View PDF

Abstract:Question Answering (QA) tasks requiring information from multiple documents often rely on a retrieval model to identify relevant information from which the reasoning model can derive an answer. The retrieval model is typically trained to maximize the likelihood of the labeled supporting evidence. However, when retrieving from large text corpora such as Wikipedia, the correct answer can often be obtained from multiple evidence candidates, not all of them labeled as positive, thus rendering the training signal weak and noisy. The problem is exacerbated when the questions are unanswerable or the answers are boolean, since the models cannot rely on lexical overlap to map answers to supporting evidences. We develop a new parameterization of set-valued retrieval that properly handles unanswerable queries, and we show that marginalizing over this set during training allows a model to mitigate false negatives in annotated supporting evidences. We test our method with two multi-document QA datasets, IIRC and HotpotQA. On IIRC, we show that joint modeling with marginalization on alternative contexts improves model performance by 5.5 F1 points and achieves a new state-of-the-art performance of 50.6 F1. We also show that marginalization results in 0.9 to 1.6 QA F1 improvement on HotpotQA in various settings.

Comments:	10 pages
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2103.12235 [cs.CL]
	(or arXiv:2103.12235v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2103.12235

Submission history

From: Ansong Ni [view email]
[v1] Mon, 22 Mar 2021 23:44:35 UTC (108 KB)
[v2] Wed, 8 Sep 2021 23:32:34 UTC (5,578 KB)

Computer Science > Computation and Language

Title:Mitigating False-Negative Contexts in Multi-document QuestionAnswering with Retrieval Marginalization

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Mitigating False-Negative Contexts in Multi-document QuestionAnswering with Retrieval Marginalization

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators