SAViR-T: Spatially Attentive Visual Reasoning with Transformers

Sahu, Pritish; Basioti, Kalliopi; Pavlovic, Vladimir

Computer Science > Computer Vision and Pattern Recognition

arXiv:2206.09265v1 (cs)

[Submitted on 18 Jun 2022 (this version), latest version 22 Jun 2022 (v2)]

Title:SAViR-T: Spatially Attentive Visual Reasoning with Transformers

Authors:Pritish Sahu, Kalliopi Basioti, Vladimir Pavlovic

View PDF

Abstract:We present a novel computational model, "SAViR-T", for the family of visual reasoning problems embodied in the Raven's Progressive Matrices (RPM). Our model considers explicit spatial semantics of visual elements within each image in the puzzle, encoded as spatio-visual tokens, and learns the intra-image as well as the inter-image token dependencies, highly relevant for the visual reasoning task. Token-wise relationship, modeled through a transformer-based SAViR-T architecture, extract group (row or column) driven representations by leveraging the group-rule coherence and use this as the inductive bias to extract the underlying rule representations in the top two row (or column) per token in the RPM. We use this relation representations to locate the correct choice image that completes the last row or column for the RPM. Extensive experiments across both synthetic RPM benchmarks, including RAVEN, I-RAVEN, RAVEN-FAIR, and PGM, and the natural image-based "V-PROM" demonstrate that SAViR-T sets a new state-of-the-art for visual reasoning, exceeding prior models' performance by a considerable margin.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2206.09265 [cs.CV]
	(or arXiv:2206.09265v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2206.09265

Submission history

From: Pritish Sahu [view email]
[v1] Sat, 18 Jun 2022 18:26:20 UTC (38,096 KB)
[v2] Wed, 22 Jun 2022 02:00:11 UTC (37,809 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:SAViR-T: Spatially Attentive Visual Reasoning with Transformers

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:SAViR-T: Spatially Attentive Visual Reasoning with Transformers

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators