How to Improve LLMs' Performance on Specific Languages: A Perspective on LLM-Derived Language Similarity
Xinhe Shi, Qingcheng Zeng, Weihao Xuan, Linchao Zhu
摘要
Large language models (LLMs) exhibit uneven performance across languages. In languagespecific applications, practitioners often rely on target-language corpora or cross-lingual transfer to achieve better performance. However, traditional linguistic typology, commonly used as a transfer language selection strategy in previous studies, may not align with LLM's perception of language similarity. This work proposes LLM-based language similarity as a novel perspective for selecting effective fine-tuning languages. We construct a framework to quantify the similarity within each language pair through both the lenses of language-specific performance patterns and cross-lingual transferability, ultimately deriving three similarity score matrices. Moreover, we observe a counterintuitive phenomenon: super-additive transfer effect, where finetuning on a certain language yields higher performance than fine-tuning directly on the target language. Additionally, due to the absence of an existing dataset meeting our experimental requirements, we construct and release the M4CQ-Pro dataset, which features domain-diverse distribution of 135 tasks and content consistency across 31 languages (including over 20 medium-and low-resource languages), with 61518 manually reviewed highquality questions per language. We evaluate our approach on representative multilingual LLMs and results show that all three LLM-based similarity measures effectively guide fine-tuning language selection, outperforming traditional linguistic similarity, with the integrated measure achieving the best results. Our approach provides not only a novel perspective on language similarity, but also practical baselines for selecting fine-tuning languages.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
相关 Paper
- Cross-Lingual Pitfalls: Automatic Probing Cross-Lingual Weakness of Multilingual Large Language ModelsZixiang Xu, Yanbo Wang, Yue Huang, Xiuying Chen 等ACL 2025 · 被引用 5 次
- Multilingual Transfer Learning for QA using Translation as Data AugmentationMihaela A. Bornea, Lin Pan, Sara Rosenthal, Radu Florian 等AAAI 2021 · 被引用 45 次
- Cross-Lingual Optimization for Language Transfer in Large Language ModelsJungseob Lee, Seongtae Hong, Hyeonseok Moon, Heuiseok LimACL 2025
- Improving the Cross-Lingual Generalisation in Visual Question AnsweringFarhad Nooralahzadeh, Rico SennrichAAAI 2023 · 被引用 8 次
- The Heterogeneous Safety Impacts of Benign Multilingual Fine-TuningWill Hawkins, Kai Rawal, Jonathan Rystrøm, Stratis Tsirtsis 等ICML 2026
