DocEmul: a Toolkit to Generate Structured Historical Documents

Capobianco, Samuele; Marinai, Simone

Computer Science > Computer Vision and Pattern Recognition

arXiv:1710.03474 (cs)

[Submitted on 10 Oct 2017]

Title:DocEmul: a Toolkit to Generate Structured Historical Documents

Authors:Samuele Capobianco, Simone Marinai

View PDF

Abstract:We propose a toolkit to generate structured synthetic documents emulating the actual document production process. Synthetic documents can be used to train systems to perform document analysis tasks. In our case we address the record counting task on handwritten structured collections containing a limited number of examples. Using the DocEmul toolkit we can generate a larger dataset to train a deep architecture to predict the number of records for each page. The toolkit is able to generate synthetic collections and also perform data augmentation to create a larger trainable dataset. It includes one method to extract the page background from real pages which can be used as a substrate where records can be written on the basis of variable structures and using cursive fonts. Moreover, it is possible to extend the synthetic collection by adding random noise, page rotations, and other visual variations. We performed some experiments on two different handwritten collections using the toolkit to generate synthetic data to train a Convolutional Neural Network able to count the number of records in the real collections.

Comments:	In Proceedings of the 14th International Conference on Document Analysis and Recognition (ICDAR), 2017
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:1710.03474 [cs.CV]
	(or arXiv:1710.03474v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1710.03474

Submission history

From: Samuele Capobianco [view email]
[v1] Tue, 10 Oct 2017 09:40:19 UTC (5,290 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:DocEmul: a Toolkit to Generate Structured Historical Documents

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:DocEmul: a Toolkit to Generate Structured Historical Documents

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators