Redwood: Using Collision Detection to Grow a Large-Scale Intent Classification Dataset

Larson, Stefan; Leach, Kevin

Computer Science > Computation and Language

arXiv:2204.05483 (cs)

[Submitted on 12 Apr 2022 (v1), last revised 25 Jul 2022 (this version, v2)]

Title:Redwood: Using Collision Detection to Grow a Large-Scale Intent Classification Dataset

Authors:Stefan Larson, Kevin Leach

View PDF

Abstract:Dialog systems must be capable of incorporating new skills via updates over time in order to reflect new use cases or deployment scenarios. Similarly, developers of such ML-driven systems need to be able to add new training data to an already-existing dataset to support these new skills. In intent classification systems, problems can arise if training data for a new skill's intent overlaps semantically with an already-existing intent. We call such cases collisions. This paper introduces the task of intent collision detection between multiple datasets for the purposes of growing a system's skillset. We introduce several methods for detecting collisions, and evaluate our methods on real datasets that exhibit collisions. To highlight the need for intent collision detection, we show that model performance suffers if new data is added in such a way that does not arbitrate colliding intents. Finally, we use collision detection to construct and benchmark a new dataset, Redwood, which is composed of 451 ntent categories from 13 original intent classification datasets, making it the largest publicly available intent classification benchmark.

Comments:	SIGDIAL 2022
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2204.05483 [cs.CL]
	(or arXiv:2204.05483v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2204.05483

Submission history

From: Stefan Larson [view email]
[v1] Tue, 12 Apr 2022 02:28:23 UTC (283 KB)
[v2] Mon, 25 Jul 2022 16:57:42 UTC (277 KB)

Computer Science > Computation and Language

Title:Redwood: Using Collision Detection to Grow a Large-Scale Intent Classification Dataset

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Redwood: Using Collision Detection to Grow a Large-Scale Intent Classification Dataset

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators