Better Simultaneous Translation with Monotonic Knowledge Distillation
Shushu Wang, Jing Wu, Kai Fan, Wei Luo, Jun Xiao, Zhongqiang Huang
Abstract
Simultaneous machine translation (SiMT) presents a unique challenge as it requires generating target tokens before the source sentence is fully consumed. This can lead to the hallucination problem, where target tokens are generated without support from the source sentence. The prefix-to-prefix training data used to train SiMT models are not always parallel, due to divergent word order between the source and target languages, and can contribute to the problem. In this paper, we propose a novel approach that leverages traditional translation models as teachers and employs a two-stage beam search algorithm to generate monotonic yet accurate reference translations for sequence-level knowledge distillation. Experimental results demonstrate the significant improvements achieved by our approach over multiple strong SiMT baselines, leading to new state-of-the-art performance across various language pairs. Notably, when evaluated on a monotonic version of the WMT15 De→En test set, which includes references generated in a more monotonic style by professional translators, our approach achieves even more substantial improvement over the baselines. The source code and data are publicly available for further exploration 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ba7b8efe-ccee-4f39-b88a-3dbb8007b05fCited by top-tier papers4
- Adapting Offline Speech Translation Models for Streaming with Future-Aware Distillation and InferenceBiao Fu, Minpeng Liao, Kai Fan, Zhongqiang Huang et al.EMNLP 2023 · 4 citations
- Adaptive Policy with Wait-k Model for Simultaneous TranslationLibo Zhao, Kai Fan, Wei Luo, Jing Wu et al.EMNLP 2023 · 2 citations
- Training Simultaneous Speech Translation with Robust and Random Wait-k-Tokens StrategyLinlin Zhang, Kai Fan, Jiajun Bu, Zhongqiang HuangEMNLP 2023 · 1 citation
- SimulPL: Aligning Human Preferences in Simultaneous Machine TranslationDonglei Yu, Yang Zhao, Jie Zhu, Yangyifan Xu et al.ICLR 2025
Builds on6
- Monotonic Multihead AttentionXutai Ma, Juan Miguel Pino, James Cross, Liezl Puzon et al.ICLR 2020 · 148 citations
- SimulSpeech: End-to-End Simultaneous Speech to Text TranslationYi Ren, Jinglin Liu, Xu Tan, Chen Zhang et al.ACL 2020 · 81 citations
- Cross Attention Augmented Transducer Networks for Simultaneous TranslationDan Liu, Mengge Du, Xiaoxi Li, Ya Li et al.EMNLP 2021 · 28 citations
- Information-Transport-based Policy for Simultaneous TranslationShaolei Zhang, Yang FengEMNLP 2022 · 25 citations
- Improving Simultaneous Machine Translation with Monolingual DataHexuan Deng, Liang Ding, Xuebo Liu, Meishan Zhang et al.AAAI 2023 · 19 citations
Related papers
- Context Consistency between Training and Inference in Simultaneous Machine TranslationMeizhi Zhong, Lemao Liu, Kehai Chen, Mingming Yang et al.ACL 2024
- Decoder-only Streaming Transformer for Simultaneous TranslationShoutao Guo, Shaolei Zhang, Yang FengACL 2024 · 3 citations
- PsFuture: A Pseudo-Future-based Zero-Shot Adaptive Policy for Simultaneous Machine TranslationLibo Zhao, Jing Li, Ziqian ZengEMNLP 2024 · 1 citation
- Self-Modifying State Modeling for Simultaneous Machine TranslationDonglei Yu, Xiaomian Kang, Yuchen Liu, Yu Zhou et al.ACL 2024
- Hidden Markov Transformer for Simultaneous Machine TranslationShaolei Zhang, Yang FengICLR 2023 · 11 citations
