Middle-Layer Representation Alignment for Cross-Lingual Transfer in Fine-Tuned LLMs
Danni Liu, Jan Niehues
摘要
While large language models demonstrate remarkable capabilities at task-specific applications through fine-tuning, extending these benefits across diverse languages is essential for broad accessibility. However, effective crosslingual transfer is hindered by LLM performance gaps across languages and the scarcity of fine-tuning data in many languages. Through analysis of LLM internal representations from over 1,000+ language pairs, we discover that middle layers exhibit the strongest potential for cross-lingual alignment. Building on this finding, we propose a middle-layer alignment objective integrated into task-specific training. Our experiments on slot filling, machine translation, and structured text generation show consistent improvements in cross-lingual transfer, especially to lower-resource languages. The method is robust to the choice of alignment languages and generalizes to languages unseen during alignment. Furthermore, we show that separately trained alignment modules can be merged with existing task-specific modules, improving cross-lingual capabilities without full re-training. Our code is publicly available 1 . 0 4 8 12 16 20 24 28 32 Layer ID 0 50 100 Avg. retrieval accuracy (%) Llama 3 0 4 8 12 16 20 24 28 Layer ID Qwen 2.5 Overall Low-res. (a) Cross-lingual semantic alignment (measured by average retrieval accuracy over 35 languages and 1190 language directions) varies by layer, with the middle layer showing the highest score. Lower-resource languages are poorly aligned. 0 20 40 60 0 25 50 75 100 Transfer result Llama 3 Correlation: 0.56 F1 0 20 40 60 Qwen 2.5 Correlation: 0.70 Cross-lingual representation retrieval accuracy (%) (b) Positive correlation between base model cross-lingual semantic alignment and downstream transfer performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Crosscoding Through Time: Tracking Emergence & Consolidation Of Linguistic Representations Throughout LLM PretrainingDeniz Bayazit, Aaron Mueller, Antoine BosselutACL 2026 · 被引用 3 次
- LinguaMap: Which Layers of LLMs Speak Your Language and How to Tune Them?J. Ben Tamo, Daniel Carlander-Reuterfelt, Jonathan Rubin, Oleg Poliannikov 等ICLR 2026 · 被引用 3 次
- SiLP: Enhancing Non-Dominant Language Capabilities with a Selective Bidirectional Language Projection FrameworkJunpeng Liu, Jiuyi Li, Kaiyu Huang, Bo Jin 等ACL 2026
- SOAPTriage: SOAP-Guided Multi-View Clinical Text Modeling Framework for Automated ESI PredictionEnming Wang, Jianlei Wang, Xueping Peng, Hongjiao Guan 等ACL 2026
- Post-Training Language Models for Crosslingual ConsistencyTianyu Liu, Jirui Qi, Mrinmaya Sachan, Ryan Cotterell 等ICML 2026
它引用的顶会 Paper25
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 被引用 741 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
相关 Paper
- LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM SafetyJunxiao Yang, Haoran Liu, Jinzhe Tu, Jiale Cheng 等ACL 2026 · 被引用 1 次
- Analyzing LLMs' Knowledge Boundary Cognition Across Languages Through the Lens of Internal RepresentationsChenghao Xiao, Hou Pong Chan, Hao Zhang, Mahani Aljunied 等ACL 2025 · 被引用 3 次
- Language on Demand, Knowledge at Core: Composing LLMs with Encoder-Decoder Translation Models for Extensible MultilingualityMengyu Bu, Yang FengACL 2026 · 被引用 2 次
- Bridging the Language Gaps in Large Language Models with Inference-Time Cross-Lingual InterventionWeixuan Wang, Minghao Wu, Barry Haddow, Alexandra BirchACL 2025 · 被引用 17 次
- AlignX: Advancing Multilingual Large Language Models with Multilingual Representation AlignmentMengyu Bu, Shaolei Zhang, Zhongjun He, Hua Wu 等EMNLP 2025
