From English to Second Language Mastery: Enhancing LLMs with Cross-Lingual Continued Instruction Tuning
Linjuan Wu, Haoran Wei, Baosong Yang, Weiming Lu
摘要
Supervised Fine-Tuning (SFT) with translated instruction data effectively adapts Large Language Models (LLMs) from English to non-English languages. We introduce Cross-Lingual Continued Instruction Tuning (X-CIT), which fully leverages translation-based parallel instruction data to enhance cross-lingual adapt-ability. X-CIT emulates the human process of second language acquisition and is guided by Chomsky’s Principles and Parameters Theory. It first fine-tunes the LLM on English instruction data to establish foundational capabilities (i.e. Principles), then continues with target language translation and customized chat-instruction data to adjust "parameters" specific to the target language. This chat-instruction data captures alignment information in translated parallel data, guiding the model to initially think and respond in its native language before transitioning to the target language. To further mimic human learning progression, we incorporate Self-Paced Learning (SPL) during continued training, allowing the model to advance from simple to complex tasks. Implemented on Llama-2-7B across five languages, X-CIT was evaluated against three objective benchmarks and an LLM-as-a-judge benchmark, improving the strongest baseline by an average of 1.97% and 8.2% in these two benchmarks, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Crosslingual Generalization through Multitask FinetuningNiklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts 等ACL 2023 · 被引用 319 次
- X-FACTR: Multilingual Factual Knowledge Retrieval from Pretrained Language ModelsZhengbao Jiang, Antonios Anastasopoulos, Jun Araki, Haibo Ding 等EMNLP 2020 · 被引用 81 次
- XCOT: Cross-lingual Instruction Tuning for Cross-lingual Chain-of-Thought ReasoningLinzheng Chai, Jian Yang, Tao Sun, Hongcheng Guo 等AAAI 2025 · 被引用 70 次
- On the Cross-lingual Transferability of Monolingual RepresentationsMikel Artetxe, Sebastian Ruder, Dani YogatamaACL 2020 · 被引用 57 次
相关 Paper
- The Heterogeneous Safety Impacts of Benign Multilingual Fine-TuningWill Hawkins, Kai Rawal, Jonathan Rystrøm, Stratis Tsirtsis 等ICML 2026
- Cross-Lingual Optimization for Language Transfer in Large Language ModelsJungseob Lee, Seongtae Hong, Hyeonseok Moon, Heuiseok LimACL 2025
- Enhancing Non-English Capabilities of English-Centric Large Language Models Through Deep Supervision Fine-TuningWenshuai Huo, Xiaocheng Feng, Yichong Huang, Chengpeng Fu 等AAAI 2025 · 被引用 11 次
- CommonIT: Commonality-Aware Instruction Tuning for Large Language Models via Data PartitionsJun Rao, Xuebo Liu, Lian Lian, Shengjun Cheng 等EMNLP 2024 · 被引用 2 次
- Self-Distillation Bridges Distribution Gap in Language Model Fine-TuningZhaorui Yang, Tianyu Pang, Haozhe Feng, Han Wang 等ACL 2024
