The Reasonableness Behind Unreasonable Translation Capability of Large Language Model
Tingchen Fu, Lemao Liu, Deng Cai, Guoping Huang, Shuming Shi, Rui Yan
摘要
Multilingual large language models (LLM) trained on non-parallel data yield impressive translation capabilities. Existing studies demonstrate that incidental sentence-level bilingualism within pre-training data contributes to the translation abilities of large language models. However, it has been observed that the translation capabilities persist even when incidental sentence-level bilingualism is excluded from the training corpus. Therefore, in this study, we comprehensively investigate the question why LLM can acquire translation capability without sentence-level bilingualism data. To this end, we examine the impacts of word-level bilingualism data (i.e., word alignment data and code-switching data) or even purified monolingual data in addition to sentence-level bilingualism data. Through extensive experiments, we have made significant findings. It turns out that word alignment data plays a crucial role in enabling LLMs to acquire translation ability. Surprisingly, the translation signal derived from word alignment data is even comparable to that obtained from sentence-level bilingualism. Moreover, pre-training on purified monolingual data may enable a slight translation signal for LLM, thanks to the shared parameters in Transformer and some shared tokens across both source and target languages. Our code is available in https://github.com/TingchenFu/ICLR24-TransContamination.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos 等USENIX Security 2019 · 被引用 1,386 次
- Large Language Models Struggle to Learn Long-Tail KnowledgeNikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace 等ICML 2023 · 被引用 623 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
相关 Paper
- The Role of Mixed-Language Documents for Multilingual Large Language Model PretrainingJiandong Shao, Raphael Tang, Crystina Zhang, Karin Sevegnani 等ACL 2026
- Just Go Parallel: Improving the Multilingual Capabilities of Large Language ModelsMuhammad Reza Qorib, Junyi Li, Hwee Tou NgACL 2025 · 被引用 5 次
- Alternating Language Modeling for Cross-Lingual Pre-TrainingJian Yang, Shuming Ma, Dongdong Zhang, Shuangzhi Wu 等AAAI 2020 · 被引用 94 次
- PreAlign: Boosting Cross-Lingual Transfer by Early Establishment of Multilingual AlignmentJiahuan Li, Shujian Huang, Aarron Ching, Xinyu Dai 等EMNLP 2024 · 被引用 5 次
- Fine-Tuning Large Language Models to Translate: Will a Touch of Noisy Data in Misaligned Languages Suffice?Dawei Zhu, Pinzhen Chen, Miaoran Zhang, Barry Haddow 等EMNLP 2024 · 被引用 3 次
