Hidden Markov Transformer for Simultaneous Machine Translation
Shaolei Zhang, Yang Feng
Abstract
Simultaneous machine translation (SiMT) outputs the target sequence while receiving the source sequence, and hence learning when to start translating each target token is the core challenge for SiMT task. However, it is non-trivial to learn the optimal moment among many possible moments of starting translating, as the moments of starting translating always hide inside the model and can only be supervised with the observed target sequence. In this paper, we propose a Hidden Markov Transformer (HMT), which treats the moments of starting translating as hidden events and the target sequence as the corresponding observed events, thereby organizing them as a hidden Markov model. HMT explicitly models multiple moments of starting translating as the candidate hidden events, and then selects one to generate the target token. During training, by maximizing the marginal likelihood of the target sequence over multiple moments of starting translating, HMT learns to start translating at the moments that target tokens can be generated more accurately. Experiments on multiple SiMT benchmarks show that HMT outperforms strong baselines and achieves state-of-the-art performance 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Unified Segment-to-Segment Framework for Simultaneous Sequence GenerationShaolei Zhang, Yang FengNeurIPS 2023 · 9 citations
- SimulMEGA: MoE Routers are Advanced Policy Makers for Simultaneous Speech TranslationChenyang Le, Bing Han, Jinshun Li, Songyong Chen et al.NeurIPS 2025 · 3 citations
- Non-autoregressive Streaming Transformer for Simultaneous TranslationZhengrui Ma, Shaolei Zhang, Shoutao Guo, Chenze Shao et al.EMNLP 2023 · 3 citations
- Large Language Models Are Read/Write Policy-Makers for Simultaneous GenerationShoutao Guo, Shaolei Zhang, Zhengrui Ma, Yang FengAAAI 2025 · 3 citations
- Adaptive Policy with Wait-k Model for Simultaneous TranslationLibo Zhao, Kai Fan, Wei Luo, Jing Wu et al.EMNLP 2023 · 2 citations
Builds on10
- Monotonic Multihead AttentionXutai Ma, Juan Miguel Pino, James Cross, Liezl Puzon et al.ICLR 2020 · 148 citations
- Future-Guided Incremental Transformer for Simultaneous TranslationShaolei Zhang, Yang Feng, Liangyou LiAAAI 2021 · 44 citations
- Cross Attention Augmented Transducer Networks for Simultaneous TranslationDan Liu, Mengge Du, Xiaoxi Li, Ya Li et al.EMNLP 2021 · 28 citations
- Modeling Dual Read/Write Paths for Simultaneous Machine TranslationShaolei Zhang, Yang FengACL 2022 · 27 citations
- Information-Transport-based Policy for Simultaneous TranslationShaolei Zhang, Yang FengEMNLP 2022 · 25 citations
Related papers
- Learning Optimal Policy for Simultaneous Machine Translation via Binary SearchShoutao Guo, Shaolei Zhang, Yang FengACL 2023 · 9 citations
- Decoder-only Streaming Transformer for Simultaneous TranslationShoutao Guo, Shaolei Zhang, Yang FengACL 2024 · 3 citations
- Translation-based Supervision for Policy Generation in Simultaneous Neural Machine TranslationAshkan Alinejad, Hassan S. Shavarani, Anoop SarkarEMNLP 2021 · 6 citations
- Self-Modifying State Modeling for Simultaneous Machine TranslationDonglei Yu, Xiaomian Kang, Yuchen Liu, Yu Zhou et al.ACL 2024
- A Generative Framework for Simultaneous Machine TranslationYishu Miao, Phil Blunsom, Lucia SpeciaEMNLP 2021 · 12 citations
