Document-Level Machine Translation with Large-Scale Public Parallel Corpora
Proyag Pal, Alexandra Birch, Kenneth Heafield
Abstract
Despite the fact that document-level machine translation has inherent advantages over sentence-level machine translation due to additional information available to a model from document context, most translation systems continue to operate at a sentence level. This is primarily due to the severe lack of publicly available large-scale parallel corpora at the document level. We release a large-scale open parallel corpus with document context extracted from ParaCrawl in five language pairs, along with code to compile document-level datasets for any language pair supported by ParaCrawl. We train context-aware models on these datasets and find improvements in terms of overall translation quality and targeted document-level phenomena. We also analyse how much long-range information is useful to model some of these discourse phenomena and find models are able to utilise context from several preceding sentences.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c814d7c3-080d-4d06-ab6f-592a843988e8Cited by top-tier papers3
- Extending Automatic Machine Translation Evaluation to Book-Length DocumentsKuang-Da Wang, Shuoyang Ding, Chao-Han Huck Yang, Ping-Chun Hsieh et al.EMNLP 2025
- AdaDPI: Document-level Translation Adaptive Agent via Dynamic Parametric InternalizationHong Ren, Liting Deng, Shaolin Zhu, Deyi XiongACL 2026
- HShare: Fast LLM Decoding by Hierarchical Key-Value SharingHuaijin Wu, Lianqiang Li, Hantao Huang, Tu Yi et al.ICLR 2025
Builds on4
- Prompting Large Language Model for Machine Translation: A Case StudyBiao Zhang, Barry Haddow, Alexandra BirchICML 2023 · 402 citations
- ParaCrawl: Web-Scale Acquisition of Parallel CorporaMarta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield et al.ACL 2020 · 132 citations
- Document-Level Machine Translation with Large Language ModelsLongyue Wang, Chenyang Lyu, Tianbo Ji, Zhirui Zhang et al.EMNLP 2023 · 129 citations
- CCAligned: A Massive Collection of Cross-Lingual Web-Document PairsAhmed El-Kishky, Vishrav Chaudhary, Francisco Guzmán, Philipp KoehnEMNLP 2020 · 6 citations
Related papers
- Discourse-Centric Evaluation of Document-level Machine Translation with a New Densely Annotated Parallel Corpus of NovelsYuchen Eleanor Jiang, Tianyu Liu, Shuming Ma, Dongdong Zhang et al.ACL 2023 · 7 citations
- Challenges in Context-Aware Neural Machine TranslationLinghao Jin, Jacqueline He, Jonathan May, Xuezhe MaEMNLP 2023 · 8 citations
- Exploring Discourse Structure in Document-level Machine TranslationXinyu Hu, Xiaojun WanEMNLP 2023 · 3 citations
- Breaking the Corpus Bottleneck for Context-Aware Neural Machine Translation with Cross-Task Pre-trainingLinqing Chen, Junhui Li, Zhengxian Gong, Boxing Chen et al.ACL 2021
- Document Graph for Neural Machine TranslationMingzhou Xu, Liangyou Li, Derek F. Wong, Qun Liu et al.EMNLP 2021
