Learning Similarity Functions for Pronunciation Variations

Naaman, Einat; Adi, Yossi; Keshet, Joseph

Computer Science > Computation and Language

arXiv:1703.09817 (cs)

[Submitted on 28 Mar 2017 (v1), last revised 18 Jun 2017 (this version, v3)]

Title:Learning Similarity Functions for Pronunciation Variations

Authors:Einat Naaman, Yossi Adi, Joseph Keshet

View PDF

Abstract:A significant source of errors in Automatic Speech Recognition (ASR) systems is due to pronunciation variations which occur in spontaneous and conversational speech. Usually ASR systems use a finite lexicon that provides one or more pronunciations for each word. In this paper, we focus on learning a similarity function between two pronunciations. The pronunciations can be the canonical and the surface pronunciations of the same word or they can be two surface pronunciations of different words. This task generalizes problems such as lexical access (the problem of learning the map** between words and their possible pronunciations), and defining word neighborhoods. It can also be used to dynamically increase the size of the pronunciation lexicon, or in predicting ASR errors. We propose two methods, which are based on recurrent neural networks, to learn the similarity function. The first is based on binary classification, and the second is based on learning the ranking of the pronunciations. We demonstrate the efficiency of our approach on the task of lexical access using a subset of the Switchboard conversational speech corpus. Results suggest that on this task our methods are superior to previous methods which are based on graphical Bayesian methods.

Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:1703.09817 [cs.CL]
	(or arXiv:1703.09817v3 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.1703.09817

Submission history

From: Joseph Keshet [view email]
[v1] Tue, 28 Mar 2017 21:47:16 UTC (208 KB)
[v2] Sun, 28 May 2017 19:05:38 UTC (209 KB)
[v3] Sun, 18 Jun 2017 11:52:54 UTC (209 KB)

Computer Science > Computation and Language

Title:Learning Similarity Functions for Pronunciation Variations

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Learning Similarity Functions for Pronunciation Variations

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators