Universal Simultaneous Machine Translation with Mixture-of-Experts Wait-k Policy
Shaolei Zhang, Yang Feng
Abstract
Simultaneous machine translation (SiMT) generates translation before reading the entire source sentence and hence it has to trade off between translation quality and latency. To fulfill the requirements of different translation quality and latency in practical applications, the previous methods usually need to train multiple SiMT models for different latency levels, resulting in large computational costs. In this paper, we propose a universal SiMT model with Mixture-of-Experts Wait-k Policy to achieve the best translation quality under arbitrary latency with only one trained model. Specifically, our method employs multi-head attention to accomplish the mixture of experts where each head is treated as a wait-k expert with its own waiting words number, and given a test latency and source inputs, the weights of the experts are accordingly adjusted to produce the best translation. Experiments on three datasets show that our method outperforms all the strong baselines under different latency, including the state-of-the-art adaptive policy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2c3d2cbd-78ae-4e8a-80ea-4be2ac942954Cited by top-tier papers11
- Hidden Markov Transformer for Simultaneous Machine TranslationShaolei Zhang, Yang FengICLR 2023 · 11 citations
- Unified Segment-to-Segment Framework for Simultaneous Sequence GenerationShaolei Zhang, Yang FengNeurIPS 2023 · 9 citations
- Learning Optimal Policy for Simultaneous Machine Translation via Binary SearchShoutao Guo, Shaolei Zhang, Yang FengACL 2023 · 9 citations
- Spatial Speech Translation: Translating Across Space With Binaural HearablesTuochao Chen, Qirui Wang, Runlin He, Shyamnath GollakotaCHI 2025 · 5 citations
- Adapting Offline Speech Translation Models for Streaming with Future-Aware Distillation and InferenceBiao Fu, Minpeng Liao, Kai Fan, Zhongqiang Huang et al.EMNLP 2023 · 4 citations
Builds on4
- Monotonic Multihead AttentionXutai Ma, Juan Miguel Pino, James Cross, Liezl Puzon et al.ICLR 2020 · 148 citations
- Future-Guided Incremental Transformer for Simultaneous TranslationShaolei Zhang, Yang Feng, Liangyou LiAAAI 2021 · 44 citations
- Learning Adaptive Segmentation Policy for Simultaneous TranslationRuiqing Zhang, Chuanqiang Zhang, Zhongjun He, Hua Wu et al.EMNLP 2020 · 41 citations
- A Mixture of h - 1 Heads is Better than h HeadsHao Peng, Roy Schwartz, Dianqi Li, Noah A. SmithACL 2020 · 25 citations
Related papers
- A Generative Framework for Simultaneous Machine TranslationYishu Miao, Phil Blunsom, Lucia SpeciaEMNLP 2021 · 12 citations
- Adaptive Policy with Wait-k Model for Simultaneous TranslationLibo Zhao, Kai Fan, Wei Luo, Jing Wu et al.EMNLP 2023 · 2 citations
- Modeling Dual Read/Write Paths for Simultaneous Machine TranslationShaolei Zhang, Yang FengACL 2022 · 27 citations
- DrFrattn: Directly Learn Adaptive Policy from Attention for Simultaneous Machine TranslationLibo Zhao, Jing Li, Ziqian ZengEMNLP 2025 · 2 citations
- Context Consistency between Training and Inference in Simultaneous Machine TranslationMeizhi Zhong, Lemao Liu, Kehai Chen, Mingming Yang et al.ACL 2024
