Beyond Sentence-Level End-to-End Speech Translation: Context Helps
Biao Zhang, Ivan Titov, Barry Haddow, Rico Sennrich
Abstract
Document-level contextual information has shown benefits to text-based machine translation, but whether and how context helps endto-end (E2E) speech translation (ST) is still under-studied. We fill this gap through extensive experiments using a simple concatenationbased context-aware ST model, paired with adaptive feature selection on speech encodings for computational efficiency. We investigate several decoding approaches, and introduce inmodel ensemble decoding which jointly performs document-and sentence-level translation using the same model. Our results on the MuST-C benchmark with Transformer demonstrate the effectiveness of context to E2E ST. Compared to sentence-level ST, context-aware ST obtains better translation quality (+0.18-2.61 BLEU), improves pronoun and homophone translation, shows better robustness to (artificial) audio segmentation errors, and reduces latency and flicker to deliver higher quality for simultaneous translation. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Revisiting End-to-End Speech-to-Text Translation From ScratchBiao Zhang, Barry Haddow, Rico SennrichICML 2022 · 46 citations
- End-to-End Single-Channel Speaker-Turn Aware Conversational Speech TranslationJuan Pablo Zuluaga-Gomez, Zhaocheng Huang, Xing Niu, Rohit Paturi et al.EMNLP 2023 · 2 citations
- Speech Sense Disambiguation: Tackling Homophone Ambiguity in End-to-End Speech TranslationTengfei Yu, Xuebo Liu, Liang Ding, Kehai Chen et al.ACL 2024
- From Simultaneous to Streaming Machine Translation by Leveraging Streaming HistoryJavier Iranzo-Sánchez, Jorge Civera, Alfons Juan-CíscarACL 2022
Builds on3
- Curriculum Pre-training for End-to-End Speech TranslationChengyi Wang, Yu Wu, Shujie Liu, Ming Zhou et al.ACL 2020 · 100 citations
- Dynamic Context Selection for Document-level Neural Machine Translation via Reinforcement LearningXiaomian Kang, Yang Zhao, Jiajun Zhang, Chengqing ZongEMNLP 2020 · 61 citations
- Direct Segmentation Models for Streaming Speech TranslationJavier Iranzo-Sánchez, Adrià Giménez-Pastor, Joan Albert Silvestre-Cerdà, Pau Baquero-Arnal et al.EMNLP 2020 · 24 citations
Related papers
- Improving End-to-End Speech Translation by Leveraging Auxiliary Speech and Text DataYuhao Zhang, Chen Xu, Bojie Hu, Chunliang Zhang et al.AAAI 2023 · 17 citations
- Learning Adaptive Segmentation Policy for End-to-End Simultaneous TranslationRuiqing Zhang, Zhongjun He, Hua Wu, Haifeng WangACL 2022 · 26 citations
- Shiftable Context: Addressing Training-Inference Context Mismatch in Simultaneous Speech TranslationMatthew Raffel, Drew Penney, Lizhong ChenICML 2023 · 4 citations
- Learning When to Translate for Streaming SpeechQian Dong, Yaoming Zhu, Mingxuan Wang, Lei LiACL 2022
- Phone Features Improve Speech TranslationElizabeth Salesky, Alan W. BlackACL 2020 · 2 citations
