Synchronous Speech Recognition and Speech-to-Text Translation with Interactive Decoding
Yuchen Liu, Jiajun Zhang, Hao Xiong, Long Zhou, Zhongjun He, Hua Wu, Haifeng Wang, Chengqing Zong
Abstract
Speech-to-text translation (ST), which translates source language speech into target language text, has attracted intensive attention in recent years. Compared to the traditional pipeline system, the end-to-end ST model has potential benefits of lower latency, smaller model size, and less error propagation. However, it is notoriously difficult to implement such a model without transcriptions as intermediate. Existing works generally apply multi-task learning to improve translation quality by jointly training end-to-end ST along with automatic speech recognition (ASR). However, different tasks in this method cannot utilize information from each other, which limits the improvement. Other works propose a two-stage model where the second model can use the hidden state from the first one, but its cascade manner greatly affects the efficiency of training and inference process. In this paper, we propose a novel interactive attention mechanism which enables ASR and ST to perform synchronously and interactively in a single model. Specifically, the generation of transcriptions and translations not only relies on its previous outputs but also the outputs predicted in the other task. Experiments on TED speech translation corpora have shown that our proposed model can outperform strong baselines on the quality of speech translation and achieve better speech recognition performances as well.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1ef5f52d-90fe-4b26-a7da-72acd8887c3aCited by top-tier papers9
- Listen, Understand and Translate: Triple Supervision Decouples End-to-end Speech-to-text TranslationQianqian Dong, Rong Ye, Mingxuan Wang, Hao Zhou et al.AAAI 2021 · 65 citations
- Other Roles Matter! Enhancing Role-Oriented Dialogue Summarization via Role InteractionsHaitao Lin, Junnan Zhu, Lu Xiang, Yu Zhou et al.ACL 2022 · 36 citations
- Regularizing End-to-End Speech Translation with Triangular Decomposition AgreementYichao Du, Zhirui Zhang, Weizhi Wang, Boxing Chen et al.AAAI 2022 · 25 citations
- Non-Parametric Domain Adaptation for End-to-End Speech TranslationYichao Du, Weizhi Wang, Zhirui Zhang, Boxing Chen et al.EMNLP 2022 · 10 citations
- Divergence-Guided Simultaneous Speech TranslationXinjie Chen, Kai Fan, Wei Luo, Linlin Zhang et al.AAAI 2024 · 6 citations
Related papers
- Consecutive Decoding for Speech-to-text TranslationQianqian Dong, Mingxuan Wang, Hao Zhou, Shuang Xu et al.AAAI 2021 · 46 citations
- Learning Adaptive Segmentation Policy for End-to-End Simultaneous TranslationRuiqing Zhang, Zhongjun He, Hua Wu, Haifeng WangACL 2022 · 26 citations
- SimulSpeech: End-to-End Simultaneous Speech to Text TranslationYi Ren, Jinglin Liu, Xu Tan, Chen Zhang et al.ACL 2020 · 81 citations
- Unified Speech-Text Pre-training for Speech Translation and RecognitionYun Tang, Hongyu Gong, Ning Dong, Changhan Wang et al.ACL 2022 · 104 citations
- SpeechUT: Bridging Speech and Text with Hidden-Unit for Encoder-Decoder Based Speech-Text Pre-trainingZiqiang Zhang, Long Zhou, Junyi Ao, Shujie Liu et al.EMNLP 2022 · 38 citations
