Showing 1–1 of 1 results for author: Gutierrez, F R

Search v0.5.6 released 2020-02-24

arXiv:2210.07978 [pdf, other]

cs.SD cs.CL eess.AS

Improving generalizability of distilled self-supervised speech processing models under distorted settings

Authors: Kuan-Po Huang, Yu-Kuan Fu, Tsu-Yuan Hsu, Fabian Ritter Gutierrez, Fan-Lin Wang, Liang-Hsuan Tseng, Yu Zhang, Hung-yi Lee

Abstract: Self-supervised learned (SSL) speech pre-trained models perform well across various speech processing tasks. Distilled versions of SSL models have been developed to match the needs of on-device speech applications. Though having similar performance as original SSL models, distilled counterparts suffer from performance degradation even more than their original versions in distorted environments. Th… ▽ More Self-supervised learned (SSL) speech pre-trained models perform well across various speech processing tasks. Distilled versions of SSL models have been developed to match the needs of on-device speech applications. Though having similar performance as original SSL models, distilled counterparts suffer from performance degradation even more than their original versions in distorted environments. This paper proposes to apply Cross-Distortion Map** and Domain Adversarial Training to SSL models during knowledge distillation to alleviate the performance gap caused by the domain mismatch problem. Results show consistent performance improvements under both in- and out-of-domain distorted setups for different downstream tasks while kee** efficient model size. △ Less

Submitted 20 October, 2022; v1 submitted 14 October, 2022; originally announced October 2022.

Comments: Accepted by IEEE SLT2022

Search v0.5.6 released 2020-02-24