Middle-Layer Representation Alignment for Cross-Lingual Transfer in Fine-Tuned LLMs
Danni Liu, Jan Niehues
Abstract
While large language models demonstrate remarkable capabilities at task-specific applications through fine-tuning, extending these benefits across diverse languages is essential for broad accessibility. However, effective crosslingual transfer is hindered by LLM performance gaps across languages and the scarcity of fine-tuning data in many languages. Through analysis of LLM internal representations from over 1,000+ language pairs, we discover that middle layers exhibit the strongest potential for cross-lingual alignment. Building on this finding, we propose a middle-layer alignment objective integrated into task-specific training. Our experiments on slot filling, machine translation, and structured text generation show consistent improvements in cross-lingual transfer, especially to lower-resource languages. The method is robust to the choice of alignment languages and generalizes to languages unseen during alignment. Furthermore, we show that separately trained alignment modules can be merged with existing task-specific modules, improving cross-lingual capabilities without full re-training. Our code is publicly available 1 . 0 4 8 12 16 20 24 28 32 Layer ID 0 50 100 Avg. retrieval accuracy (%) Llama 3 0 4 8 12 16 20 24 28 Layer ID Qwen 2.5 Overall Low-res. (a) Cross-lingual semantic alignment (measured by average retrieval accuracy over 35 languages and 1190 language directions) varies by layer, with the middle layer showing the highest score. Lower-resource languages are poorly aligned. 0 20 40 60 0 25 50 75 100 Transfer result Llama 3 Correlation: 0.56 F1 0 20 40 60 Qwen 2.5 Correlation: 0.70 Cross-lingual representation retrieval accuracy (%) (b) Positive correlation between base model cross-lingual semantic alignment and downstream transfer performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f5a06bd-7ce6-4d25-8a9f-02f9b3f0a088Cited by top-tier papers5
- Crosscoding Through Time: Tracking Emergence & Consolidation Of Linguistic Representations Throughout LLM PretrainingDeniz Bayazit, Aaron Mueller, Antoine BosselutACL 2026 · 3 citations
- LinguaMap: Which Layers of LLMs Speak Your Language and How to Tune Them?J. Ben Tamo, Daniel Carlander-Reuterfelt, Jonathan Rubin, Oleg Poliannikov et al.ICLR 2026 · 3 citations
- SiLP: Enhancing Non-Dominant Language Capabilities with a Selective Bidirectional Language Projection FrameworkJunpeng Liu, Jiuyi Li, Kaiyu Huang, Bo Jin et al.ACL 2026
- SOAPTriage: SOAP-Guided Multi-View Clinical Text Modeling Framework for Automated ESI PredictionEnming Wang, Jianlei Wang, Xueping Peng, Hongjiao Guan et al.ACL 2026
- Post-Training Language Models for Crosslingual ConsistencyTianyu Liu, Jirui Qi, Mrinmaya Sachan, Ryan Cotterell et al.ICML 2026
Builds on25
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 741 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
Related papers
- LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM SafetyJunxiao Yang, Haoran Liu, Jinzhe Tu, Jiale Cheng et al.ACL 2026 · 1 citation
- Analyzing LLMs' Knowledge Boundary Cognition Across Languages Through the Lens of Internal RepresentationsChenghao Xiao, Hou Pong Chan, Hao Zhang, Mahani Aljunied et al.ACL 2025 · 3 citations
- Language on Demand, Knowledge at Core: Composing LLMs with Encoder-Decoder Translation Models for Extensible MultilingualityMengyu Bu, Yang FengACL 2026 · 2 citations
- Bridging the Language Gaps in Large Language Models with Inference-Time Cross-Lingual InterventionWeixuan Wang, Minghao Wu, Barry Haddow, Alexandra BirchACL 2025 · 17 citations
- AlignX: Advancing Multilingual Large Language Models with Multilingual Representation AlignmentMengyu Bu, Shaolei Zhang, Zhongjun He, Hua Wu et al.EMNLP 2025
