Iterative Shallow Fusion of Backward Language Model for End-to-End Speech Recognition

Ogawa, Atsunori; Moriya, Takafumi; Kamo, Naoyuki; Tawara, Naohiro; Delcroix, Marc

Electrical Engineering and Systems Science > Audio and Speech Processing

arXiv:2310.11010 (eess)

[Submitted on 17 Oct 2023]

Title:Iterative Shallow Fusion of Backward Language Model for End-to-End Speech Recognition

Authors:Atsunori Ogawa, Takafumi Moriya, Naoyuki Kamo, Naohiro Tawara, Marc Delcroix

View PDF

Abstract:We propose a new shallow fusion (SF) method to exploit an external backward language model (BLM) for end-to-end automatic speech recognition (ASR). The BLM has complementary characteristics with a forward language model (FLM), and the effectiveness of their combination has been confirmed by rescoring ASR hypotheses as post-processing. In the proposed SF, we iteratively apply the BLM to partial ASR hypotheses in the backward direction (i.e., from the possible next token to the start symbol) during decoding, substituting the newly calculated BLM scores for the scores calculated at the last iteration. To enhance the effectiveness of this iterative SF (ISF), we train a partial sentence-aware BLM (PBLM) using reversed text data including partial sentences, considering the framework of ISF. In experiments using an attention-based encoder-decoder ASR system, we confirmed that ISF using the PBLM shows comparable performance with SF using the FLM. By performing ISF, early pruning of prospective hypotheses can be prevented during decoding, and we can obtain a performance improvement compared to applying the PBLM as post-processing. Finally, we confirmed that, by combining SF and ISF, further performance improvement can be obtained thanks to the complementarity of the FLM and PBLM.

Comments:	Accepted to ICASSP 2023
Subjects:	Audio and Speech Processing (eess.AS); Computation and Language (cs.CL)
Cite as:	arXiv:2310.11010 [eess.AS]
	(or arXiv:2310.11010v1 [eess.AS] for this version)
	https://doi.org/10.48550/arXiv.2310.11010

Submission history

From: Atsunori Ogawa [view email]
[v1] Tue, 17 Oct 2023 05:44:10 UTC (103 KB)

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Iterative Shallow Fusion of Backward Language Model for End-to-End Speech Recognition

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Audio and Speech Processing

Title:Iterative Shallow Fusion of Backward Language Model for End-to-End Speech Recognition

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators