Discriminative Sampling of Proposals in Self-Supervised Transformers for Weakly Supervised Object Localization

Murtaza, Shakeeb; Belharbi, Soufiane; Pedersoli, Marco; Sarraf, Aydin; Granger, Eric

Computer Science > Computer Vision and Pattern Recognition

arXiv:2209.09209 (cs)

[Submitted on 9 Sep 2022 (v1), last revised 20 Nov 2022 (this version, v2)]

Title:Discriminative Sampling of Proposals in Self-Supervised Transformers for Weakly Supervised Object Localization

Authors:Shakeeb Murtaza, Soufiane Belharbi, Marco Pedersoli, Aydin Sarraf, Eric Granger

View PDF

Abstract:Drones are employed in a growing number of visual recognition applications. A recent development in cell tower inspection is drone-based asset surveillance, where the autonomous flight of a drone is guided by localizing objects of interest in successive aerial images. In this paper, we propose a method to train deep weakly-supervised object localization (WSOL) models based only on image-class labels to locate object with high confidence. To train our localizer, pseudo labels are efficiently harvested from a self-supervised vision transformers (SSTs). However, since SSTs decompose the scene into multiple maps containing various object parts, and do not rely on any explicit supervisory signal, they cannot distinguish between the object of interest and other objects, as required WSOL. To address this issue, we propose leveraging the multiple maps generated by the different transformer heads to acquire pseudo-labels for training a deep WSOL model. In particular, a new Discriminative Proposals Sampling (DiPS) method is introduced that relies on a CNN classifier to identify discriminative regions. Then, foreground and background pixels are sampled from these regions in order to train a WSOL model for generating activation maps that can accurately localize objects belonging to a specific class. Empirical results on the challenging TelDrone dataset indicate that our proposed approach can outperform state-of-art methods over a wide range of threshold values over produced maps. We also computed results on CUB dataset, showing that our method can be adapted for other tasks.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2209.09209 [cs.CV]
	(or arXiv:2209.09209v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2209.09209

Submission history

From: Shakeeb Murtaza [view email]
[v1] Fri, 9 Sep 2022 18:33:23 UTC (47,844 KB)
[v2] Sun, 20 Nov 2022 02:55:22 UTC (7,431 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Discriminative Sampling of Proposals in Self-Supervised Transformers for Weakly Supervised Object Localization

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Discriminative Sampling of Proposals in Self-Supervised Transformers for Weakly Supervised Object Localization

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators