Hidden Markov Transformer for Simultaneous Machine Translation
Shaolei Zhang, Yang Feng
摘要
Simultaneous machine translation (SiMT) outputs the target sequence while receiving the source sequence, and hence learning when to start translating each target token is the core challenge for SiMT task. However, it is non-trivial to learn the optimal moment among many possible moments of starting translating, as the moments of starting translating always hide inside the model and can only be supervised with the observed target sequence. In this paper, we propose a Hidden Markov Transformer (HMT), which treats the moments of starting translating as hidden events and the target sequence as the corresponding observed events, thereby organizing them as a hidden Markov model. HMT explicitly models multiple moments of starting translating as the candidate hidden events, and then selects one to generate the target token. During training, by maximizing the marginal likelihood of the target sequence over multiple moments of starting translating, HMT learns to start translating at the moments that target tokens can be generated more accurately. Experiments on multiple SiMT benchmarks show that HMT outperforms strong baselines and achieves state-of-the-art performance 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Unified Segment-to-Segment Framework for Simultaneous Sequence GenerationShaolei Zhang, Yang FengNeurIPS 2023 · 被引用 9 次
- SimulMEGA: MoE Routers are Advanced Policy Makers for Simultaneous Speech TranslationChenyang Le, Bing Han, Jinshun Li, Songyong Chen 等NeurIPS 2025 · 被引用 3 次
- Non-autoregressive Streaming Transformer for Simultaneous TranslationZhengrui Ma, Shaolei Zhang, Shoutao Guo, Chenze Shao 等EMNLP 2023 · 被引用 3 次
- Large Language Models Are Read/Write Policy-Makers for Simultaneous GenerationShoutao Guo, Shaolei Zhang, Zhengrui Ma, Yang FengAAAI 2025 · 被引用 3 次
- Adaptive Policy with Wait-k Model for Simultaneous TranslationLibo Zhao, Kai Fan, Wei Luo, Jing Wu 等EMNLP 2023 · 被引用 2 次
它引用的顶会 Paper10
- Monotonic Multihead AttentionXutai Ma, Juan Miguel Pino, James Cross, Liezl Puzon 等ICLR 2020 · 被引用 148 次
- Future-Guided Incremental Transformer for Simultaneous TranslationShaolei Zhang, Yang Feng, Liangyou LiAAAI 2021 · 被引用 44 次
- Cross Attention Augmented Transducer Networks for Simultaneous TranslationDan Liu, Mengge Du, Xiaoxi Li, Ya Li 等EMNLP 2021 · 被引用 28 次
- Modeling Dual Read/Write Paths for Simultaneous Machine TranslationShaolei Zhang, Yang FengACL 2022 · 被引用 27 次
- Information-Transport-based Policy for Simultaneous TranslationShaolei Zhang, Yang FengEMNLP 2022 · 被引用 25 次
相关 Paper
- Learning Optimal Policy for Simultaneous Machine Translation via Binary SearchShoutao Guo, Shaolei Zhang, Yang FengACL 2023 · 被引用 9 次
- Decoder-only Streaming Transformer for Simultaneous TranslationShoutao Guo, Shaolei Zhang, Yang FengACL 2024 · 被引用 3 次
- Translation-based Supervision for Policy Generation in Simultaneous Neural Machine TranslationAshkan Alinejad, Hassan S. Shavarani, Anoop SarkarEMNLP 2021 · 被引用 6 次
- Self-Modifying State Modeling for Simultaneous Machine TranslationDonglei Yu, Xiaomian Kang, Yuchen Liu, Yu Zhou 等ACL 2024
- A Generative Framework for Simultaneous Machine TranslationYishu Miao, Phil Blunsom, Lucia SpeciaEMNLP 2021 · 被引用 12 次
