Exploring the GLIDE model for Human Action-effect Prediction

Li, Fangjun; Hogg, David C.; Cohn, Anthony G.

Computer Science > Computer Vision and Pattern Recognition

arXiv:2208.01136 (cs)

[Submitted on 1 Aug 2022]

Title:Exploring the GLIDE model for Human Action-effect Prediction

Authors:Fangjun Li, David C. Hogg, Anthony G. Cohn

View PDF

Abstract:We address the following action-effect prediction task. Given an image depicting an initial state of the world and an action expressed in text, predict an image depicting the state of the world following the action. The prediction should have the same scene context as the input image. We explore the use of the recently proposed GLIDE model for performing this task. GLIDE is a generative neural network that can synthesize (inpaint) masked areas of an image, conditioned on a short piece of text. Our idea is to mask-out a region of the input image where the effect of the action is expected to occur. GLIDE is then used to inpaint the masked region conditioned on the required action. In this way, the resulting image has the same background context as the input image, updated to show the effect of the action. We give qualitative results from experiments using the EPIC dataset of ego-centric videos labelled with actions.

Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2208.01136 [cs.CV]
	(or arXiv:2208.01136v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2208.01136

Submission history

From: Fangjun Li [view email]
[v1] Mon, 1 Aug 2022 20:51:39 UTC (14,050 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Exploring the GLIDE model for Human Action-effect Prediction

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Exploring the GLIDE model for Human Action-effect Prediction

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators