MaCLR: Motion-aware Contrastive Learning of Representations for Videos

Xiao, Fanyi; Tighe, Joseph; Modolo, Davide

Computer Science > Computer Vision and Pattern Recognition

arXiv:2106.09703 (cs)

[Submitted on 17 Jun 2021 (v1), last revised 20 Jul 2022 (this version, v2)]

Title:MaCLR: Motion-aware Contrastive Learning of Representations for Videos

Authors:Fanyi Xiao, Joseph Tighe, Davide Modolo

View PDF

Abstract:We present MaCLR, a novel method to explicitly perform cross-modal self-supervised video representations learning from visual and motion modalities. Compared to previous video representation learning methods that mostly focus on learning motion cues implicitly from RGB inputs, MaCLR enriches standard contrastive learning objectives for RGB video clips with a cross-modal learning objective between a Motion pathway and a Visual pathway. We show that the representation learned with our MaCLR method focuses more on foreground motion regions and thus generalizes better to downstream tasks. To demonstrate this, we evaluate MaCLR on five datasets for both action recognition and action detection, and demonstrate state-of-the-art self-supervised performance on all datasets. Furthermore, we show that MaCLR representation can be as effective as representations learned with full supervision on UCF101 and HMDB51 action recognition, and even outperform the supervised representation for action recognition on VidSitu and SSv2, and action detection on AVA.

Comments:	ECCV 2022
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2106.09703 [cs.CV]
	(or arXiv:2106.09703v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2106.09703

Submission history

From: Davide Modolo [view email]
[v1] Thu, 17 Jun 2021 17:57:11 UTC (2,374 KB)
[v2] Wed, 20 Jul 2022 16:38:23 UTC (524 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.CV

< prev | next >

new | recent | 2021-06

Change to browse by:

References & Citations

DBLP - CS Bibliography

listing | bibtex

Fanyi Xiao
Davide Modolo

export BibTeX citation

Computer Science > Computer Vision and Pattern Recognition

Title:MaCLR: Motion-aware Contrastive Learning of Representations for Videos

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:MaCLR: Motion-aware Contrastive Learning of Representations for Videos

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators