Diverse Pretrained Context Encodings Improve Document Translation
Domenic Donato, Lei Yu, Chris Dyer
Abstract
We propose a new architecture for adapting a sentence-level sequence-to-sequence transformer by incorporating multiple pretrained document context signals and assess the impact on translation performance of (1) different pretraining approaches for generating these signals, (2) the quantity of parallel data for which document context is available, and (3) conditioning on source, target, or source and target contexts. Experiments on the NIST Chinese-English, and IWSLT and WMT English-German tasks support four general conclusions: that using pretrained context representations markedly improves sample efficiency, that adequate parallel data resources are crucial for learning to use document context, that jointly conditioning on multiple context representations outperforms any single representation, and that source context is more valuable for translation performance than target side context. Our best multicontext model consistently outperforms the best existing context-aware transformers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2bcdaf49-254e-40a7-9830-e73fa822102cBuilds on3
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- Incorporating BERT into Neural Machine TranslationJinhua Zhu, Yingce Xia, Lijun Wu, Di He et al.ICLR 2020 · 391 citations
Related papers
- Breaking the Corpus Bottleneck for Context-Aware Neural Machine Translation with Cross-Task Pre-trainingLinqing Chen, Junhui Li, Zhengxian Gong, Boxing Chen et al.ACL 2021
- Multilingual Document-Level Translation Enables Zero-Shot Transfer From Sentences to DocumentsBiao Zhang, Ankur Bapna, Melvin Johnson, Ali Dabirmoghaddam et al.ACL 2022 · 15 citations
- Divide and Rule: Effective Pre-Training for Context-Aware Multi-Encoder Translation ModelsLorenzo Lupo, Marco Dinarelli, Laurent BesacierACL 2022 · 21 citations
- Dynamic Context Selection for Document-level Neural Machine Translation via Reinforcement LearningXiaomian Kang, Yang Zhao, Jiajun Zhang, Chengqing ZongEMNLP 2020 · 61 citations
- Document Graph for Neural Machine TranslationMingzhou Xu, Liangyou Li, Derek F. Wong, Qun Liu et al.EMNLP 2021
