Language-based Action Concept Spaces Improve Video Self-Supervised Learning

Ranasinghe, Kanchana; Ryoo, Michael

Computer Science > Computer Vision and Pattern Recognition

arXiv:2307.10922 (cs)

[Submitted on 20 Jul 2023 (v1), last revised 26 Oct 2023 (this version, v3)]

Title:Language-based Action Concept Spaces Improve Video Self-Supervised Learning

Authors:Kanchana Ranasinghe, Michael Ryoo

View PDF

Abstract:Recent contrastive language image pre-training has led to learning highly transferable and robust image representations. However, adapting these models to video domains with minimal supervision remains an open problem. We explore a simple step in that direction, using language tied self-supervised learning to adapt an image CLIP model to the video domain. A backbone modified for temporal modeling is trained under self-distillation settings with train objectives operating in an action concept space. Feature vectors of various action concepts extracted from a language encoder using relevant textual prompts construct this space. We introduce two train objectives, concept distillation and concept alignment, that retain generality of original representations while enforcing relations between actions and their attributes. Our approach improves zero-shot and linear probing performance on three action recognition benchmarks.

Comments:	Presented at NeurIPS 2023
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
Cite as:	arXiv:2307.10922 [cs.CV]
	(or arXiv:2307.10922v3 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2307.10922

Submission history

From: Kanchana Ranasinghe [view email]
[v1] Thu, 20 Jul 2023 14:47:50 UTC (300 KB)
[v2] Mon, 2 Oct 2023 12:57:16 UTC (301 KB)
[v3] Thu, 26 Oct 2023 14:34:55 UTC (170 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Language-based Action Concept Spaces Improve Video Self-Supervised Learning

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Language-based Action Concept Spaces Improve Video Self-Supervised Learning

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators