DocumentNet: Bridging the Data Gap in Document Pre-Training

Yu, Lijun; Miao, **; Sun, ** entity spaces from different datasets hinder the knowledge transfer between document types. In this paper, we propose a method to collect massive-scale and weakly labeled data from the web to benefit the training of VDER models. The collected dataset, named DocumentNet, does not depend on specific document types or entity sets, making it universally applicable to all VDER tasks. The current DocumentNet consists of 30M documents spanning nearly 400 document types organized in a four-level ontology. Experiments on a set of broadly adopted VDER tasks show significant improvements when DocumentNet is incorporated into the pre-training for both classic and few-shot learning settings. With the recent emergence of large language models (LLMs), DocumentNet provides a large data source to extend their multi-modal capabilities for VDER.

Computer Science > Computation and Language

arXiv:2306.08937v3 (cs)

[Submitted on 15 Jun 2023 (v1), last revised 26 Oct 2023 (this version, v3)]

Title:DocumentNet: Bridging the Data Gap in Document Pre-Training

Authors:Lijun Yu, ** Miao, Xiaoyu Sun, Jiayi Chen, Alexander G. Hauptmann, Hanjun Dai, Wei Wei

View PDF

Abstract:Document understanding tasks, in particular, Visually-rich Document Entity Retrieval (VDER), have gained significant attention in recent years thanks to their broad applications in enterprise AI. However, publicly available data have been scarce for these tasks due to strict privacy constraints and high annotation costs. To make things worse, the non-overlap** entity spaces from different datasets hinder the knowledge transfer between document types. In this paper, we propose a method to collect massive-scale and weakly labeled data from the web to benefit the training of VDER models. The collected dataset, named DocumentNet, does not depend on specific document types or entity sets, making it universally applicable to all VDER tasks. The current DocumentNet consists of 30M documents spanning nearly 400 document types organized in a four-level ontology. Experiments on a set of broadly adopted VDER tasks show significant improvements when DocumentNet is incorporated into the pre-training for both classic and few-shot learning settings. With the recent emergence of large language models (LLMs), DocumentNet provides a large data source to extend their multi-modal capabilities for VDER.

Comments:	EMNLP 2023
Subjects:	Computation and Language (cs.CL); Information Retrieval (cs.IR)
Cite as:	arXiv:2306.08937 [cs.CL]
	(or arXiv:2306.08937v3 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2306.08937

Submission history

From: Lijun Yu [view email]
[v1] Thu, 15 Jun 2023 08:21:15 UTC (1,312 KB)
[v2] Tue, 10 Oct 2023 13:48:12 UTC (1,766 KB)
[v3] Thu, 26 Oct 2023 16:23:15 UTC (1,766 KB)

Computer Science > Computation and Language

Title:DocumentNet: Bridging the Data Gap in Document Pre-Training

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:DocumentNet: Bridging the Data Gap in Document Pre-Training

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators