Integration of Text-maps in Convolutional Neural Networks for Region Detection among Different Textual Categories

Arroyo, Roberto; Tovar, Javier; Delgado, Francisco J.; Almazán, Emilio J.; Serrador, Diego G.; Hurtado, Antonio

Computer Science > Computer Vision and Pattern Recognition

arXiv:1905.10858 (cs)

[Submitted on 26 May 2019]

Title:Integration of Text-maps in Convolutional Neural Networks for Region Detection among Different Textual Categories

Authors:Roberto Arroyo, Javier Tovar, Francisco J. Delgado, Emilio J. Almazán, Diego G. Serrador, Antonio Hurtado

View PDF

Abstract:In this work, we propose a new technique that combines appearance and text in a Convolutional Neural Network (CNN), with the aim of detecting regions of different textual categories. We define a novel visual representation of the semantic meaning of text that allows a seamless integration in a standard CNN architecture. This representation, referred to as text-map, is integrated with the actual image to provide a much richer input to the network. Text-maps are colored with different intensities depending on the relevance of the words recognized over the image. Concretely, these words are previously extracted using Optical Character Recognition (OCR) and they are colored according to the probability of belonging to a textual category of interest. In this sense, this solution is especially relevant in the context of item coding for supermarket products, where different types of textual categories must be identified, such as ingredients or nutritional facts. We evaluated our solution in the proprietary item coding dataset of Nielsen Brandbank, which contains more than 10,000 images for train and 2,000 images for test. The reported results demonstrate that our approach focused on visual and textual data outperforms state-of-the-art algorithms only based on appearance, such as standard Faster R-CNN. These enhancements are reflected in precision and recall, which are improved in 42 and 33 points respectively.

Comments:	Conference on Computer Vision and Pattern Recognition (CVPR). Language and Vision Workshop 2019
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:1905.10858 [cs.CV]
	(or arXiv:1905.10858v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1905.10858

Submission history

From: Roberto Arroyo [view email]
[v1] Sun, 26 May 2019 18:59:32 UTC (5,401 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Integration of Text-maps in Convolutional Neural Networks for Region Detection among Different Textual Categories

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Integration of Text-maps in Convolutional Neural Networks for Region Detection among Different Textual Categories

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators