Mask-Align: Self-Supervised Neural Word Alignment
Chi Chen, Maosong Sun, Yang Liu
摘要
Word alignment, which aims to align translationally equivalent words between source and target sentences, plays an important role in many natural language processing tasks. Current unsupervised neural alignment methods focus on inducing alignments from neural machine translation models, which does not leverage the full context in the target sequence. In this paper, we propose MASK-ALIGN, a selfsupervised word alignment model that takes advantage of the full context on the target side. Our model parallelly masks out each target token and predicts it conditioned on both source and the remaining target tokens. This two-step process is based on the assumption that the source token contributing most to recovering the masked target token should be aligned. We also introduce an attention variant called leaky attention, which alleviates the problem of high cross-attention weights on specific tokens such as periods. Experiments on four language pairs show that our model outperforms previous unsupervised neural aligners and obtains new state-of-the-art results. 1 * Corresponding author 1 Code can be found at https://github.com/THUNLP-MT/ Mask-Align .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- What learning algorithm is in-context learning? Investigations with linear modelsEkin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma 等ICLR 2023 · 被引用 85 次
- Prompting Neural Machine Translation with Translation MemoriesAbudurexiti Reheman, Tao Zhou, Yingfeng Luo, Di Yang 等AAAI 2023 · 被引用 11 次
- Cross-Align: Modeling Deep Cross-lingual Interactions for Word AlignmentSiyu Lai, Zhen Yang, Fandong Meng, Yufeng Chen 等EMNLP 2022 · 被引用 6 次
- Self-Supervised Quality Estimation for Machine TranslationYuanhang Zheng, Zhixing Tan, Meng Zhang, Mieradilijiang Maimaiti 等EMNLP 2021 · 被引用 5 次
- BinaryAlign: Word Alignment as Binary Sequence LabelingGaetan Latouche, Marc-André Carbonneau, Benjamin SwansonACL 2024 · 被引用 2 次
它引用的顶会 Paper3
- Non-autoregressive Machine Translation with Disentangled Context TransformerJungo Kasai, James Cross, Marjan Ghazvininejad, Jiatao GuICML 2020 · 被引用 113 次
- Accurate Word Alignment Induction from Neural Machine TranslationYun Chen, Yang Liu, Guanhua Chen, Xin Jiang 等EMNLP 2020 · 被引用 56 次
- End-to-End Neural Word Alignment Outperforms GIZA++Thomas Zenkel, Joern Wuebker, John DeNeroACL 2020 · 被引用 2 次
相关 Paper
- A Bidirectional Transformer Based Alignment Model for Unsupervised Word AlignmentJingyi Zhang, Josef van GenabithACL 2021
- Self-supervised Bilingual Syntactic Alignment for Neural Machine TranslationTianfu Zhang, Heyan Huang, Chong Feng, Longbing CaoAAAI 2021 · 被引用 7 次
- Improving Pretrained Cross-Lingual Language Models via Self-Labeled Word AlignmentZewen Chi, Li Dong, Bo Zheng, Shaohan Huang 等ACL 2021
- A Supervised Word Alignment Method based on Cross-Language Span Prediction using Multilingual BERTMasaaki Nagata, Katsuki Chousa, Masaaki NishinoEMNLP 2020 · 被引用 38 次
- SenseBERT: Driving Some Sense into BERTYoav Levine, Barak Lenz, Or Dagan, Ori Ram 等ACL 2020 · 被引用 27 次
