G-Transformer for Document-Level Machine Translation
Guangsheng Bao, Yue Zhang, Zhiyang Teng, Boxing Chen, Weihua Luo
摘要
Document-level MT models are still far from satisfactory. Existing work extend translation unit from single sentence to multiple sentences. However, study shows that when we further enlarge the translation unit to a whole document, supervised training of Transformer can fail. In this paper, we find such failure is not caused by overfitting, but by sticking around local minima during training. Our analysis shows that the increased complexity of target-to-source attention is a reason for the failure. As a solution, we propose G-Transformer, introducing locality assumption as an inductive bias into Transformer, reducing the hypothesis space of the attention from target to source. Experiments show that G-Transformer converges faster and more stably than Transformer, achieving new state-of-the-art BLEU scores for both nonpretraining and pre-training settings on three benchmark datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Document-Level Machine Translation with Large Language ModelsLongyue Wang, Chenyang Lyu, Tianbo Ji, Zhirui Zhang 等EMNLP 2023 · 被引用 129 次
- Target-Side Augmentation for Document-Level Machine TranslationGuangsheng Bao, Zhiyang Teng, Yue ZhangACL 2023 · 被引用 9 次
- DeMPT: Decoding-enhanced Multi-phase Prompt Tuning for Making LLMs Be Better Context-aware TranslatorsXinglin Lyu, Junhui Li, Yanqing Zhao, Min Zhang 等EMNLP 2024 · 被引用 4 次
- Exploring Discourse Structure in Document-level Machine TranslationXinyu Hu, Xiaojun WanEMNLP 2023 · 被引用 3 次
- Modeling Consistency Preference via Lexical Chains for Document-level Neural Machine TranslationXinglin Lyu, Junhui Li, Shimin Tao, Hao Yang 等EMNLP 2022 · 被引用 3 次
它引用的顶会 Paper3
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Contextualized Rewriting for Text SummarizationGuangsheng Bao, Yue ZhangAAAI 2021 · 被引用 17 次
相关 Paper
- Diverse Pretrained Context Encodings Improve Document TranslationDomenic Donato, Lei Yu, Chris DyerACL 2021
- ETC: Encoding Long and Structured Inputs in TransformersJoshua Ainslie, Santiago Ontañón, Chris Alberti, Vaclav Cvicek 等EMNLP 2020 · 被引用 268 次
- G2LFormer: Global-to-Local Query Enhancement for Robust Table Structure RecognitionHaosheng Cai, Yang XueACM MM 2025
- Span Graph Transformer for Document-Level Named Entity RecognitionHongli Mao, Xian-Ling Mao, Hanlin Tang, Yuming Shang 等AAAI 2024 · 被引用 3 次
- Hidden Markov Transformer for Simultaneous Machine TranslationShaolei Zhang, Yang FengICLR 2023 · 被引用 11 次
