Breaking the Corpus Bottleneck for Context-Aware Neural Machine Translation with Cross-Task Pre-training
Linqing Chen, Junhui Li, Zhengxian Gong, Boxing Chen, Weihua Luo, Min Zhang, Guodong Zhou
Abstract
Context-aware neural machine translation (NMT) remains challenging due to the lack of large-scale document-level parallel dataset. To break the corpus bottleneck, in this paper we aim to improve context-aware NMT by taking the advantage of the availability of both large-scale sentence-level parallel dataset and source-side monolingual documents. 1 To this end, we propose two pre-training tasks. One learns to translate a sentence from source language to target language on the sentencelevel parallel dataset while the other learns to translate a document from deliberately noised to original on the monolingual documents. Importantly, the two pre-training tasks are jointly and simultaneously learned via the same model, thereafter fine-tuned on scalelimited parallel documents from both sentencelevel and document-level perspectives. Experimental results on four translation tasks show that our approach significantly improves translation performance. One nice property of our approach is that the fine-tuned model can be used to translate both sentences and documents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5f72841d-c6a8-40e4-8069-0573b0c3f7b2Builds on2
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- Dynamic Context Selection for Document-level Neural Machine Translation via Reinforcement LearningXiaomian Kang, Yang Zhao, Jiajun Zhang, Chengqing ZongEMNLP 2020 · 61 citations
Related papers
- Diverse Pretrained Context Encodings Improve Document TranslationDomenic Donato, Lei Yu, Chris DyerACL 2021
- Divide and Rule: Effective Pre-Training for Context-Aware Multi-Encoder Translation ModelsLorenzo Lupo, Marco Dinarelli, Laurent BesacierACL 2022 · 21 citations
- Challenges in Context-Aware Neural Machine TranslationLinghao Jin, Jacqueline He, Jonathan May, Xuezhe MaEMNLP 2023 · 8 citations
- MT2: Towards a Multi-Task Machine Translation Model with Translation-Specific In-Context LearningChunyou Li, Mingtong Liu, Hongxiao Zhang, Yufeng Chen et al.EMNLP 2023 · 3 citations
- Acquiring Knowledge from Pre-Trained Model to Neural Machine TranslationRongxiang Weng, Heng Yu, Shujian Huang, Shanbo Cheng et al.AAAI 2020 · 71 citations
