Integrating Scene Text and Visual Appearance for Fine-Grained Image Classification with Convolutional Neural Networks

Bai, Xiang; Yang, Mingkun; Lyu, Pengyuan; Xu, Yongchao

Computer Science > Computer Vision and Pattern Recognition

arXiv:1704.04613v1 (cs)

[Submitted on 15 Apr 2017 (this version), latest version 30 May 2017 (v2)]

Title:Integrating Scene Text and Visual Appearance for Fine-Grained Image Classification with Convolutional Neural Networks

Authors:Xiang Bai, Mingkun Yang, Pengyuan Lyu, Yongchao Xu

View PDF

Abstract:Text in natural images contains plenty of semantics that are often highly relevant to objects or scene. In this paper, we are concerned with the problem on fully exploiting scene text for visual understanding. The basic idea is combining word representations and deep visual features into a globally trainable deep convolutional neural network. First, the recognized words are obtained by a scene text reading system. Then, we combine the word embedding of the recognized words and the deep visual features into a single representation, which is optimized by a convolutional neural network for fine-grained image classification. In our framework, the attention mechanism is adopted to reveal the relevance between each recognized word and the given image, which reasonably enhances the recognition performance. We have performed experiments on two datasets: Con-Text dataset and Drink Bottle dataset, that are proposed for fine-grained classification of business places and drink bottles respectively. The experimental results consistently demonstrate that the proposed method combining textual and visual cues significantly outperforms classification with only visual representations. Moreover, we have shown that the learned representation improves the retrieval performance on the drink bottle images by a large margin, which is potentially useful in product search.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:1704.04613 [cs.CV]
	(or arXiv:1704.04613v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1704.04613

Submission history

From: Yang Mingkun [view email]
[v1] Sat, 15 Apr 2017 09:44:08 UTC (4,939 KB)
[v2] Tue, 30 May 2017 01:27:20 UTC (4,940 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Integrating Scene Text and Visual Appearance for Fine-Grained Image Classification with Convolutional Neural Networks

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Integrating Scene Text and Visual Appearance for Fine-Grained Image Classification with Convolutional Neural Networks

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators