Dense Image Representation with Spatial Pyramid VLAD Coding of CNN for Locally Robust Captioning

Shin, Andrew; Yamaguchi, Masataka; Ohnishi, Katsunori; Harada, Tatsuya

Computer Science > Computer Vision and Pattern Recognition

arXiv:1603.09046 (cs)

[Submitted on 30 Mar 2016]

Title:Dense Image Representation with Spatial Pyramid VLAD Coding of CNN for Locally Robust Captioning

Authors:Andrew Shin, Masataka Yamaguchi, Katsunori Ohnishi, Tatsuya Harada

View PDF

Abstract:The workflow of extracting features from images using convolutional neural networks (CNN) and generating captions with recurrent neural networks (RNN) has become a de-facto standard for image captioning task. However, since CNN features are originally designed for classification task, it is mostly concerned with the main conspicuous element of the image, and often fails to correctly convey information on local, secondary elements. We propose to incorporate coding with vector of locally aggregated descriptors (VLAD) on spatial pyramid for CNN features of sub-regions in order to generate image representations that better reflect the local information of the images. Our results show that our method of compact VLAD coding can match CNN features with as little as 3% of dimensionality and, when combined with spatial pyramid, it results in image captions that more accurately take local elements into account.

Comments:	submitted to ECCV2016
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:1603.09046 [cs.CV]
	(or arXiv:1603.09046v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1603.09046

Submission history

From: Andrew Shin [view email]
[v1] Wed, 30 Mar 2016 05:48:05 UTC (9,080 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.CV

< prev | next >

new | recent | 2016-03

Change to browse by:

References & Citations

DBLP - CS Bibliography

listing | bibtex

Andrew Shin
Masataka Yamaguchi
Katsunori Ohnishi
Tatsuya Harada

export BibTeX citation

Computer Science > Computer Vision and Pattern Recognition

Title:Dense Image Representation with Spatial Pyramid VLAD Coding of CNN for Locally Robust Captioning

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Dense Image Representation with Spatial Pyramid VLAD Coding of CNN for Locally Robust Captioning

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators