Scene Summarization: Clustering Scene Videos into Spatially Diverse Frames

Chen, Chao; Zhu, Mingzhi; Singh, Ankush Pratap; Yan, Yu; Xu, Felix Juefei; Feng, Chen

Computer Science > Computer Vision and Pattern Recognition

arXiv:2311.17940 (cs)

[Submitted on 28 Nov 2023]

Title:Scene Summarization: Clustering Scene Videos into Spatially Diverse Frames

Authors:Chao Chen, Mingzhi Zhu, Ankush Pratap Singh, Yu Yan, Felix Juefei Xu, Chen Feng

View PDF

Abstract:We propose scene summarization as a new video-based scene understanding task. It aims to summarize a long video walkthrough of a scene into a small set of frames that are spatially diverse in the scene, which has many impotant applications, such as in surveillance, real estate, and robotics. It stems from video summarization but focuses on long and continuous videos from moving cameras, instead of user-edited fragmented video clips that are more commonly studied in existing video summarization works. Our solution to this task is a two-stage self-supervised pipeline named SceneSum. Its first stage uses clustering to segment the video sequence. Our key idea is to combine visual place recognition (VPR) into this clustering process to promote spatial diversity. Its second stage needs to select a representative keyframe from each cluster as the summary while respecting resource constraints such as memory and disk space limits. Additionally, if the ground truth image trajectory is available, our method can be easily augmented with a supervised loss to enhance the clustering and keyframe selection. Extensive experiments on both real-world and simulated datasets show our method outperforms common video summarization baselines by 50%

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2311.17940 [cs.CV]
	(or arXiv:2311.17940v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2311.17940

Submission history

From: Chao Chen [view email]
[v1] Tue, 28 Nov 2023 22:18:26 UTC (10,555 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Scene Summarization: Clustering Scene Videos into Spatially Diverse Frames

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Scene Summarization: Clustering Scene Videos into Spatially Diverse Frames

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators