Accurate Word Alignment Induction from Neural Machine Translation
Yun Chen, Yang Liu, Guanhua Chen, Xin Jiang, Qun Liu
摘要
Despite its original goal to jointly learn to align and translate, prior researches suggest that Transformer captures poor word alignments through its attention mechanism. In this paper, we show that attention weights DO capture accurate word alignments and propose two novel word alignment induction methods SHIFT-ATT and SHIFT-AET. The main idea is to induce alignments at the step when the to-be-aligned target token is the decoder input rather than the decoder output as in previous work. SHIFT-ATT is an interpretation method that induces alignments from the attention weights of Transformer and does not require parameter update or architecture change. SHIFT-AET extracts alignments from an additional alignment module which is tightly integrated into Transformer and trained in isolation with supervision from symmetrized SHIFT-ATT alignments. Experiments on three publicly available datasets demonstrate that both methods perform better than their corresponding neural baselines and SHIFT-AET significantly outperforms GIZA++ by 1.4-4.8 AER points. 1 * Corresponding author. Part of the work was done when Yun was in Huawei Noah's Ark Lab. 1 Code can be found at https://github.com/ sufe-nlp/transformer-alignment . das weiß ich . das weiß ich . das weiß ich .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Sparse Attention with Linear UnitsBiao Zhang, Ivan Titov, Rico SennrichEMNLP 2021 · 被引用 32 次
- Lexically Constrained Neural Machine Translation with Explicit Alignment GuidanceGuanhua Chen, Yun Chen, Victor O. K. LiAAAI 2021 · 被引用 29 次
- Information-Transport-based Policy for Simultaneous TranslationShaolei Zhang, Yang FengEMNLP 2022 · 被引用 25 次
- Towards Opening the Black Box of Neural Machine Translation: Source and Target Interpretations of the TransformerJavier Ferrando, Gerard I. Gállego, Belen Alastruey, Carlos Escolano 等EMNLP 2022 · 被引用 18 次
- Generative Retrieval as Multi-Vector Dense RetrievalShiguang Wu, Wenda Wei, Mengqi Zhang, Zhumin Chen 等SIGIR 2024 · 被引用 14 次
它引用的顶会 Paper3
- On Identifiability in TransformersGino Brunner, Yang Liu, Damian Pascual, Oliver Richter 等ICLR 2020 · 被引用 210 次
- Alignment-Enhanced Transformer for Constraining NMT with Pre-Specified TranslationsKai Song, Kun Wang, Heng Yu, Yue Zhang 等AAAI 2020 · 被引用 49 次
- End-to-End Neural Word Alignment Outperforms GIZA++Thomas Zenkel, Joern Wuebker, John DeNeroACL 2020 · 被引用 2 次
相关 Paper
- A Bidirectional Transformer Based Alignment Model for Unsupervised Word AlignmentJingyi Zhang, Josef van GenabithACL 2021
- Mask-Align: Self-Supervised Neural Word AlignmentChi Chen, Maosong Sun, Yang LiuACL 2021
- Cross-Align: Modeling Deep Cross-lingual Interactions for Word AlignmentSiyu Lai, Zhen Yang, Fandong Meng, Yufeng Chen 等EMNLP 2022 · 被引用 6 次
- Attention is Not Only a Weight: Analyzing Transformers with Vector NormsGoro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro InuiEMNLP 2020 · 被引用 138 次
- Aligner-Encoders: Self-Attention Transformers Can Be Self-TransducersAdam Stooke, Rohit Prabhavalkar, Khe Chai Sim, Pedro Moreno MengibarNeurIPS 2024 · 被引用 4 次
