Monotonic Multihead Attention
Xutai Ma, Juan Miguel Pino, James Cross, Liezl Puzon, Jiatao Gu
摘要
Simultaneous machine translation models start generating a target sequence before they have encoded or read the source sequence. Recent approaches for this task either apply a fixed policy on a state-of-the art Transformer model, or a learnable monotonic attention on a weaker recurrent neural network-based structure. In this paper, we propose a new attention mechanism, Monotonic Multihead Attention (MMA), which extends the monotonic attention mechanism to multihead attention. We also introduce two novel and interpretable approaches for latency control that are specifically designed for multiple attentions heads. We apply MMA to the simultaneous machine translation task and demonstrate better latency-quality tradeoffs compared to MILk, the previous state-of-the-art approach. We also analyze how the latency controls affect the attention span and we motivate the introduction of our model by analyzing the effect of the number of decoder layers and heads on quality and latency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- Fast Transformers with Clustered AttentionApoorv Vyas, Angelos Katharopoulos, François FleuretNeurIPS 2020 · 被引用 193 次
- Future-Guided Incremental Transformer for Simultaneous TranslationShaolei Zhang, Yang Feng, Liangyou LiAAAI 2021 · 被引用 44 次
- Learning Adaptive Segmentation Policy for Simultaneous TranslationRuiqing Zhang, Chuanqiang Zhang, Zhongjun He, Hua Wu 等EMNLP 2020 · 被引用 41 次
- SimulSLT: End-to-End Simultaneous Sign Language TranslationAoxiong Yin, Zhou Zhao, Jinglin Liu, Weike Jin 等ACM MM 2021 · 被引用 35 次
- Modeling Dual Read/Write Paths for Simultaneous Machine TranslationShaolei Zhang, Yang FengACL 2022 · 被引用 27 次
相关 Paper
- Universal Simultaneous Machine Translation with Mixture-of-Experts Wait-k PolicyShaolei Zhang, Yang FengEMNLP 2021 · 被引用 19 次
- DrFrattn: Directly Learn Adaptive Policy from Attention for Simultaneous Machine TranslationLibo Zhao, Jing Li, Ziqian ZengEMNLP 2025 · 被引用 2 次
- Cross Attention Augmented Transducer Networks for Simultaneous TranslationDan Liu, Mengge Du, Xiaoxi Li, Ya Li 等EMNLP 2021 · 被引用 28 次
- A Generative Framework for Simultaneous Machine TranslationYishu Miao, Phil Blunsom, Lucia SpeciaEMNLP 2021 · 被引用 12 次
- Non-autoregressive Streaming Transformer for Simultaneous TranslationZhengrui Ma, Shaolei Zhang, Shoutao Guo, Chenze Shao 等EMNLP 2023 · 被引用 3 次
