Speaker conditioned acoustic modeling for multi-speaker conversational ASR

Chetupalli, Srikanth Raj; Ganapathy, Sriram

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:2104.01882 (eess)

[Submitted on 5 Apr 2021 (v1), last revised 29 Aug 2022 (this version, v2)]

Title:Speaker conditioned acoustic modeling for multi-speaker conversational ASR

Authors:Srikanth Raj Chetupalli, Sriram Ganapathy

View PDF

Abstract:In this paper, we propose a novel approach for the transcription of speech conversations with natural speaker overlap, from single channel speech recordings. The proposed model is a combination of a speaker diarization system and a hybrid automatic speech recognition (ASR) system. The speaker conditioned acoustic model (SCAM) in the ASR system consists of a series of embedding layers which use the speaker activity inputs from the diarization system to derive speaker specific embeddings. The output of the SCAM are speaker specific senones that are used for decoding the transcripts for each speaker in the conversation. In this work, we experiment with the automatic speaker activity decisions generated using an end-to-end speaker diarization system. A joint learning approach is also proposed where the diarization model and the ASR acoustic model are jointly optimized. The experiments are performed on the mixed-channel two speaker recordings from the Switchboard corpus of telephone conversations. In these experiments, we show that the proposed acoustic model, incorporating speaker activity decisions and joint optimization, improves significantly over the ASR system with explicit source filtering (relative improvements of 12% in word error rate (WER) over the baseline system).

Comments:	Manuscript accepted for presentation at Interspeech 2022
Subjects:	Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2104.01882 [eess.AS]
	(or arXiv:2104.01882v2 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.2104.01882

Submission history

From: Srikanth Raj Chetupalli [view email]
[v1] Mon, 5 Apr 2021 12:41:53 UTC (109 KB)
[v2] Mon, 29 Aug 2022 10:34:27 UTC (399 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Speaker conditioned acoustic modeling for multi-speaker conversational ASR

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Speaker conditioned acoustic modeling for multi-speaker conversational ASR

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators