$FPDM$: Domain-Specific Fast Pre-training Technique using Document-Level Metadata

Nandy, Abhilash; Kapadnis, Manav Nitin; Patnaik, Sohan; Butala, Yash Parag; Goyal, Pawan; Ganguly, Niloy

Computer Science > Computation and Language

arXiv:2306.06190v1 (cs)

[Submitted on 9 Jun 2023 (this version), latest version 14 Nov 2023 (v2)]

Title:$FPDM$: Domain-Specific Fast Pre-training Technique using Document-Level Metadata

Authors:Abhilash Nandy, Manav Nitin Kapadnis, Sohan Patnaik, Yash Parag Butala, Pawan Goyal, Niloy Ganguly

View PDF

Abstract:Pre-training Transformers has shown promising results on open-domain and domain-specific downstream tasks. However, state-of-the-art Transformers require an unreasonably large amount of pre-training data and compute. In this paper, we propose $FPDM$ (Fast Pre-training Technique using Document Level Metadata), a novel, compute-efficient framework that utilizes Document metadata and Domain-Specific Taxonomy as supervision signals to pre-train transformer encoder on a domain-specific corpus. The main innovation is that during domain-specific pretraining, an open-domain encoder is continually pre-trained using sentence-level embeddings as inputs (to accommodate long documents), however, fine-tuning is done with token-level embeddings as inputs to this encoder. We show that $FPDM$ outperforms several transformer-based baselines in terms of character-level F1 scores and other automated metrics in the Customer Support, Scientific, and Legal Domains, and shows a negligible drop in performance on open-domain benchmarks. Importantly, the novel use of document-level supervision along with sentence-level embedding input for pre-training reduces pre-training compute by around $1,000$, $4,500$, and $500$ times compared to MLM and/or NSP in Customer Support, Scientific, and Legal Domains, respectively. Code and datasets are available at this https URL.

Comments:	23 pages, 7 figures
Subjects:	Computation and Language (cs.CL); Machine Learning (cs.LG)
MSC classes:	68T50
ACM classes:	I.2.7
Cite as:	arXiv:2306.06190 [cs.CL]
	(or arXiv:2306.06190v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2306.06190

Submission history

From: Abhilash Nandy [view email]
[v1] Fri, 9 Jun 2023 18:42:19 UTC (2,092 KB)
[v2] Tue, 14 Nov 2023 21:51:21 UTC (4,321 KB)

Computer Science > Computation and Language

Title:$FPDM$: Domain-Specific Fast Pre-training Technique using Document-Level Metadata

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:$FPDM$: Domain-Specific Fast Pre-training Technique using Document-Level Metadata

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators