Human-like Controllable Image Captioning with Verb-specific Semantic Roles

Chen, Long; Jiang, Zhihong; Xiao, Jun; Liu, Wei

Computer Science > Computer Vision and Pattern Recognition

arXiv:2103.12204 (cs)

[Submitted on 22 Mar 2021]

Title:Human-like Controllable Image Captioning with Verb-specific Semantic Roles

Authors:Long Chen, Zhihong Jiang, Jun Xiao, Wei Liu

View PDF

Abstract:Controllable Image Captioning (CIC) -- generating image descriptions following designated control signals -- has received unprecedented attention over the last few years. To emulate the human ability in controlling caption generation, current CIC studies focus exclusively on control signals concerning objective properties, such as contents of interest or descriptive patterns. However, we argue that almost all existing objective control signals have overlooked two indispensable characteristics of an ideal control signal: 1) Event-compatible: all visual contents referred to in a single sentence should be compatible with the described activity. 2) Sample-suitable: the control signals should be suitable for a specific image sample. To this end, we propose a new control signal for CIC: Verb-specific Semantic Roles (VSR). VSR consists of a verb and some semantic roles, which represents a targeted activity and the roles of entities involved in this activity. Given a designated VSR, we first train a grounded semantic role labeling (GSRL) model to identify and ground all entities for each role. Then, we propose a semantic structure planner (SSP) to learn human-like descriptive semantic structures. Lastly, we use a role-shift captioning model to generate the captions. Extensive experiments and ablations demonstrate that our framework can achieve better controllability than several strong baselines on two challenging CIC benchmarks. Besides, we can generate multi-level diverse captions easily. The code is available at: this https URL.

Comments:	Accepted by CVPR 2021. The code is available at: this https URL
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Multimedia (cs.MM)
Cite as:	arXiv:2103.12204 [cs.CV]
	(or arXiv:2103.12204v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2103.12204

Submission history

From: Long Chen [view email]
[v1] Mon, 22 Mar 2021 22:17:42 UTC (4,222 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Human-like Controllable Image Captioning with Verb-specific Semantic Roles

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Human-like Controllable Image Captioning with Verb-specific Semantic Roles

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators