Semi-supervised acoustic and language model training for English-isiZulu code-switched speech recognition

Biswas, A.; de Wet, F.; van der Westhuizen, E.; Niesler, T. R.

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:2004.04054 (eess)

[Submitted on 5 Apr 2020]

Title:Semi-supervised acoustic and language model training for English-isiZulu code-switched speech recognition

Authors:A. Biswas, F. de Wet, E. van der Westhuizen, T.R. Niesler

View PDF

Abstract:We present an analysis of semi-supervised acoustic and language model training for English-isiZulu code-switched ASR using soap opera speech. Approximately 11 hours of untranscribed multilingual speech was transcribed automatically using four bilingual code-switching transcription systems operating in English-isiZulu, English-isiXhosa, English-Setswana and English-Sesotho. These transcriptions were incorporated into the acoustic and language model training sets. Results showed that the TDNN-F acoustic models benefit from the additional semi-supervised data and that even better performance could be achieved by including additional CNN layers. Using these CNN-TDNN-F acoustic models, a first iteration of semi-supervised training achieved an absolute mixed-language WER reduction of 3.4%, and a further 2.2% after a second iteration. Although the languages in the untranscribed data were unknown, the best results were obtained when all automatically transcribed data was used for training and not just the utterances classified as English-isiZulu. Despite reducing perplexity, the semi-supervised language model was not able to improve the ASR performance.

Comments:	4th Code-Switch workshop, France
Subjects:	Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD); Machine Learning (stat.ML)
Cite as:	arXiv:2004.04054 [eess.AS]
	(or arXiv:2004.04054v1 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.2004.04054

Submission history

From: Astik Biswas [view email]
[v1] Sun, 5 Apr 2020 06:27:29 UTC (3,884 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Semi-supervised acoustic and language model training for English-isiZulu code-switched speech recognition

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Semi-supervised acoustic and language model training for English-isiZulu code-switched speech recognition

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators