학술논문

Multi-task Language Modeling for Improving Speech Recognition of Rare Words

Document Type

Working Paper

Author

Yang, Chao-Han Huck; Liu, Linda; Gandhe, Ankur; Gu, Yile; Raju, Anirudh; Filimonov, Denis; Bulyko, Ivan

Source

Subject

Computer Science - Computation and Language
Computer Science - Artificial Intelligence
Computer Science - Machine Learning
Computer Science - Neural and Evolutionary Computing
Computer Science - Sound
Electrical Engineering and Systems Science - Audio and Speech Processing

Language

Abstract

End-to-end automatic speech recognition (ASR) systems are increasingly popular due to their relative architectural simplicity and competitive performance. However, even though the average accuracy of these systems may be high, the performance on rare content words often lags behind hybrid ASR systems. To address this problem, second-pass rescoring is often applied leveraging upon language modeling. In this paper, we propose a second-pass system with multi-task learning, utilizing semantic targets (such as intent and slot prediction) to improve speech recognition performance. We show that our rescoring model trained with these additional tasks outperforms the baseline rescoring model, trained with only the language modeling task, by 1.4% on a general test and by 2.6% on a rare word test set in terms of word-error-rate relative (WERR). Our best ASR system with multi-task LM shows 4.6% WERR deduction compared with RNN Transducer only ASR baseline for rare words recognition.
Comment: Accepted to IEEE Automatic Speech Recognition and Understanding (ASRU) 2021

Online Access

Open Access (Arxiv) Find it@PNU

이메일

부산대학교 도서관

Online Access

메일 발송