Artificial Disfluency Detection, Uh No, Disfluency Generation for the Masses

Passali, T.; Mavropoulos, T.; Tsoumakas, G.; Meditskos, G.; Vrochidis, S.

Computer Science > Computation and Language

arXiv:2211.09235v1 (cs)

[Submitted on 16 Nov 2022]

Title:Artificial Disfluency Detection, Uh No, Disfluency Generation for the Masses

Authors:T. Passali, T. Mavropoulos, G. Tsoumakas, G. Meditskos, S. Vrochidis

View PDF

Abstract:Existing approaches for disfluency detection typically require the existence of large annotated datasets. However, current datasets for this task are limited, suffer from class imbalance, and lack some types of disfluencies that can be encountered in real-world scenarios. This work proposes LARD, a method for automatically generating artificial disfluencies from fluent text. LARD can simulate all the different types of disfluencies (repetitions, replacements and restarts) based on the reparandum/interregnum annotation scheme. In addition, it incorporates contextual embeddings into the disfluency generation to produce realistic context-aware artificial disfluencies. Since the proposed method requires only fluent text, it can be used directly for training, bypassing the requirement of annotated disfluent data. Our empirical evaluation demonstrates that LARD can indeed be effectively used when no or only a few data are available. Furthermore, our detailed analysis suggests that the proposed method generates realistic disfluencies and increases the accuracy of existing disfluency detectors.

Comments:	10 pages
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2211.09235 [cs.CL]
	(or arXiv:2211.09235v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2211.09235

Submission history

From: Tatiana Passali [view email]
[v1] Wed, 16 Nov 2022 22:00:02 UTC (578 KB)

Computer Science > Computation and Language

Title:Artificial Disfluency Detection, Uh No, Disfluency Generation for the Masses

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Artificial Disfluency Detection, Uh No, Disfluency Generation for the Masses

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators