What You Say Is What You Show: Visual Narration Detection in Instructional Videos

Ashutosh, Kumar; Girdhar, Rohit; Torresani, Lorenzo; Grauman, Kristen

Computer Science > Computer Vision and Pattern Recognition

arXiv:2301.02307v1 (cs)

[Submitted on 5 Jan 2023 (this version), latest version 18 Jul 2023 (v2)]

Title:What You Say Is What You Show: Visual Narration Detection in Instructional Videos

Authors:Kumar Ashutosh, Rohit Girdhar, Lorenzo Torresani, Kristen Grauman

View PDF

Abstract:Narrated "how-to" videos have emerged as a promising data source for a wide range of learning problems, from learning visual representations to training robot policies. However, this data is extremely noisy, as the narrations do not always describe the actions demonstrated in the video. To address this problem we introduce the novel task of visual narration detection, which entails determining whether a narration is visually depicted by the actions in the video. We propose "What You Say is What You Show" (WYS^2), a method that leverages multi-modal cues and pseudo-labeling to learn to detect visual narrations with only weakly labeled data. We further generalize our approach to operate on only audio input, learning properties of the narrator's voice that hint if they are currently doing what they describe. Our model successfully detects visual narrations in in-the-wild videos, outperforming strong baselines, and we demonstrate its impact for state-of-the-art summarization and alignment of instructional video.

Comments:	Technical Report
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2301.02307 [cs.CV]
	(or arXiv:2301.02307v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2301.02307

Submission history

From: Kumar Ashutosh [view email]
[v1] Thu, 5 Jan 2023 21:43:19 UTC (29,434 KB)
[v2] Tue, 18 Jul 2023 17:29:16 UTC (7,151 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:What You Say Is What You Show: Visual Narration Detection in Instructional Videos

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:What You Say Is What You Show: Visual Narration Detection in Instructional Videos

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators