Token Alignment Heads: Unveiling Attention's Role in LLM Multilingual Translation
Binbin Liu, Wenhan Han, Feng Chen, Yifan Zhang, Ping Guo, Haobin Lin, Bingni Zhang, Taifeng Wang, Yin Zheng
Abstract
Recently, large language models (LLMs) have made remarkable progress, with multilingual capability emerging as a core foundational strengths. However, the internal mechanisms by which these models perform translation remain incompletely understood. In this paper, we elucidate the relationship between the attention mechanism in LLMs and their translation abilities. We find that certain attention heads, which we term token alignment heads, are specifically responsible for mapping tokens from the source language to the target language during inference. Through a systematic investigation across various models, we confirm that these token alignment heads exhibit several key characteristics: (1) Universality: They are present in all LLMs we studied. (2) Sparsity: They constitute only a small fraction of all attention heads. (3) Consistency: The set of token alignment heads activated by the model shows strong consistency across different language pairs. (4) Causality: Interventionally removing these heads leads to a sharp decline in the model's translation performance, while randomly removing non-token alignment heads has little impact on translation ability. (5) Functional Specificity: Ablating token alignment heads disproportionately harms translation but has a varied impact on other multilingual tasks. We also traced the formation of token alignment heads during pre-training, revealing an evolutionary path of rapid proliferation, stabilization, and eventual pruning. Furthermore we leverage these token alignment heads to filter multilingual training data, and our experiments show that these data could enhance translation capabilities of the models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d027e84-0c49-4ddc-a8dd-d459e84f5bddCited by top-tier papers1
Ask how each one uses itBuilds on15
- Function Vectors in Large Language ModelsEric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller et al.ICLR 2024 · 229 citations
- How do Large Language Models Handle Multilingualism?Yiran Zhao, Wenxuan Zhang, Guizhen Chen, Kenji Kawaguchi et al.NeurIPS 2024 · 196 citations
- On the Cross-lingual Transferability of Monolingual RepresentationsMikel Artetxe, Sebastian Ruder, Dani YogatamaACL 2020 · 57 citations
- The Geometry of Multilingual Language Model RepresentationsTyler A. Chang, Zhuowen Tu, Benjamin K. BergenEMNLP 2022 · 22 citations
- Large Language Models Only Pass Primary School Exams in Indonesia: A Comprehensive Test on IndoMMLUFajri Koto, Nurul Aisyah, Haonan Li, Timothy BaldwinEMNLP 2023 · 11 citations
Related papers
- Focusing on Language: Revealing and Exploiting Language Attention Heads in Multilingual Large Language ModelsXin Liu, Qiyang Song, Qihang Zhou, Haichao Du et al.AAAI 2026
- The Same but Different: Structural Similarities and Differences in Multilingual Language ModelingRuochen Zhang, Qinan Yu, Matianyu Zang, Carsten Eickhoff et al.ICLR 2025
- TokAlign: Efficient Vocabulary Adaptation via Token AlignmentChong Li, Jiajun Zhang, Chengqing ZongACL 2025 · 7 citations
- Losing Heads in the Lottery: Pruning Transformer Attention in Neural Machine TranslationMaximiliana Behnke, Kenneth HeafieldEMNLP 2020 · 50 citations
- Exploring the Translation Mechanism of Large Language ModelsHongbin Zhang, Kehai Chen, Xuefeng Bai, Xiucheng Li et al.NeurIPS 2025 · 4 citations
