Exploring Intrinsic Language-specific Subspaces in Fine-tuning Multilingual Neural Machine Translation
Zhe Cao, Zhi Qu, Hidetaka Kamigaito, Taro Watanabe
Abstract
Multilingual neural machine translation models support fine-tuning hundreds of languages simultaneously. However, fine-tuning on full parameters solely is inefficient potentially leading to negative interactions among languages. In this work, we demonstrate that the fine-tuning for a language occurs in its intrinsic languagespecific subspace with a tiny fraction of entire parameters. Thus, we propose languagespecific LoRA to isolate intrinsic languagespecific subspaces. Furthermore, we propose architecture learning techniques and introduce a gradual pruning schedule during fine-tuning to exhaustively explore the optimal setting and the minimal intrinsic subspaces for each language, resulting in a lightweight yet effective fine-tuning procedure. The experimental results on a 12-language subset and a 30language subset of FLORES-101 show that our methods not only outperform full-parameter fine-tuning up to 2.25 spBLEU scores but also reduce trainable parameters to 0.4% for high and medium-resource languages and 1.6% for low-resource ones. Code will be released at https://github.com/Spike0924/LSLo .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Registering Source Tokens to Target Language Spaces in Multilingual Neural Machine TranslationZhi Qu, Yiran Wang, Jiannan Mao, Chenchen Ding et al.ACL 2025
- MLAS-LoRA: Language-Aware Parameters Detection and LoRA-Based Knowledge Transfer for Multilingual Machine TranslationTianyu Dong, Bo Li, Jinsong Liu, Shaolin Zhu et al.ACL 2025
Builds on12
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- VeRA: Vector-based Random Matrix AdaptationDawid Jan Kopiczko, Tijmen Blankevoort, Yuki M. AsanoICLR 2024 · 308 citations
- Transformer Feed-Forward Layers Are Key-Value MemoriesMor Geva, Roei Schuster, Jonathan Berant, Omer LevyEMNLP 2021 · 33 citations
- Parameter Differentiation Based Multilingual Neural Machine TranslationQian Wang, Jiajun ZhangAAAI 2022 · 21 citations
- Towards Higher Pareto Frontier in Multilingual Machine TranslationYi-Chong Huang, Xiaocheng Feng, Xinwei Geng, Baohang Li et al.ACL 2023 · 8 citations
Related papers
- Gradient-based Gradual Pruning for Language-Specific Multilingual Neural Machine TranslationDan He, Minh-Quang Pham, Thanh-Le Ha, Marco TurchiEMNLP 2023 · 2 citations
- Batched Low-Rank Adaptation of Foundation ModelsYeming Wen, Swarat ChaudhuriICLR 2024 · 32 citations
- MTL-LoRA: Low-Rank Adaptation for Multi-Task LearningYaming Yang, Dilxat Muhtar, Yelong Shen, Yuefeng Zhan et al.AAAI 2025 · 23 citations
- Positional Cognitive Specialization: Where Do LLMs Learn to Comprehend and Speak Your Language?Luis Frentzen Salim, Lun-Wei Ku, Hsing-Kuo Kenneth PaoAAAI 2026
- Learning Language Specific Sub-network for Multilingual Machine TranslationZehui Lin, Liwei Wu, Mingxuan Wang, Lei LiACL 2021
