The Reasonableness Behind Unreasonable Translation Capability of Large Language Model
Tingchen Fu, Lemao Liu, Deng Cai, Guoping Huang, Shuming Shi, Rui Yan
Abstract
Multilingual large language models (LLM) trained on non-parallel data yield impressive translation capabilities. Existing studies demonstrate that incidental sentence-level bilingualism within pre-training data contributes to the translation abilities of large language models. However, it has been observed that the translation capabilities persist even when incidental sentence-level bilingualism is excluded from the training corpus. Therefore, in this study, we comprehensively investigate the question why LLM can acquire translation capability without sentence-level bilingualism data. To this end, we examine the impacts of word-level bilingualism data (i.e., word alignment data and code-switching data) or even purified monolingual data in addition to sentence-level bilingualism data. Through extensive experiments, we have made significant findings. It turns out that word alignment data plays a crucial role in enabling LLMs to acquire translation ability. Surprisingly, the translation signal derived from word alignment data is even comparable to that obtained from sentence-level bilingualism. Moreover, pre-training on purified monolingual data may enable a slight translation signal for LLM, thanks to the shared parameters in Transformer and some shared tokens across both source and target languages. Our code is available in https://github.com/TingchenFu/ICLR24-TransContamination.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5144cc56-8f27-4fc2-8dda-84e9475e81ddCited by top-tier papers1
Ask how each one uses itBuilds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos et al.USENIX Security 2019 · 1,386 citations
- Large Language Models Struggle to Learn Long-Tail KnowledgeNikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace et al.ICML 2023 · 623 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
Related papers
- The Role of Mixed-Language Documents for Multilingual Large Language Model PretrainingJiandong Shao, Raphael Tang, Crystina Zhang, Karin Sevegnani et al.ACL 2026
- Just Go Parallel: Improving the Multilingual Capabilities of Large Language ModelsMuhammad Reza Qorib, Junyi Li, Hwee Tou NgACL 2025 · 5 citations
- Alternating Language Modeling for Cross-Lingual Pre-TrainingJian Yang, Shuming Ma, Dongdong Zhang, Shuangzhi Wu et al.AAAI 2020 · 94 citations
- PreAlign: Boosting Cross-Lingual Transfer by Early Establishment of Multilingual AlignmentJiahuan Li, Shujian Huang, Aarron Ching, Xinyu Dai et al.EMNLP 2024 · 5 citations
- Fine-Tuning Large Language Models to Translate: Will a Touch of Noisy Data in Misaligned Languages Suffice?Dawei Zhu, Pinzhen Chen, Miaoran Zhang, Barry Haddow et al.EMNLP 2024 · 3 citations
