Are we asking the right questions in MovieQA?

Jasani, Bhavan; Girdhar, Rohit; Ramanan, Deva

Computer Science > Computer Vision and Pattern Recognition

arXiv:1911.03083 (cs)

[Submitted on 8 Nov 2019]

Title:Are we asking the right questions in MovieQA?

Authors:Bhavan Jasani, Rohit Girdhar, Deva Ramanan

View PDF

Abstract:Joint vision and language tasks like visual question answering are fascinating because they explore high-level understanding, but at the same time, can be more prone to language biases. In this paper, we explore the biases in the MovieQA dataset and propose a strikingly simple model which can exploit them. We find that using the right word embedding is of utmost importance. By using an appropriately trained word embedding, about half the Question-Answers (QAs) can be answered by looking at the questions and answers alone, completely ignoring narrative context from video clips, subtitles, and movie scripts. Compared to the best published papers on the leaderboard, our simple question + answer only model improves accuracy by 5% for video + subtitle category, 5% for subtitle, 15% for DVS and 6% higher for scripts.

Comments:	Spotlight presentation at CLVL workshop, ICCV 2019. Project page: this https URL
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
Cite as:	arXiv:1911.03083 [cs.CV]
	(or arXiv:1911.03083v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1911.03083

Submission history

From: Bhavan Jasani [view email]
[v1] Fri, 8 Nov 2019 06:49:45 UTC (4,124 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.CL

< prev | next >

new | recent | 2019-11

Change to browse by:

cs
cs.CV

References & Citations

DBLP - CS Bibliography

listing | bibtex

Bhavan Jasani
Rohit Girdhar
Deva Ramanan

export BibTeX citation

Computer Science > Computer Vision and Pattern Recognition

Title:Are we asking the right questions in MovieQA?

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Are we asking the right questions in MovieQA?

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators