Finding the Translation Switch: Discovering and Exploiting the Task-Initiation Features in LLMs
Xinwei Wu, Heng Liu, Xiaohu Zhao, Yuqi Ren, Linlong Xu, Longyue Wang, Deyi Xiong, Weihua Luo, Kaifu Zhang
Abstract
Large Language Models (LLMs) frequently exhibit strong translation abilities, even without task-specific fine-tuning. However, the internal mechanisms governing this innate capability remain largely opaque. To demystify this process, we leverage Sparse Autoencoders (SAEs) and introduce a novel framework for identifying task-specific features. Our method first recalls features that are frequently co-activated on translation inputs and then filters them for functional coherence using a PCA-based consistency metric. This framework successfully isolates a small set of "translation initiation" features. Causal interventions demonstrate that amplifying these features steers the model towards correct translation, while ablating them induces hallucinations and off-task outputs, confirming they represent a core component of the model's innate translation competency. Moving from analysis to application, we leverage this mechanistic insight to propose a new data selection strategy for efficient fine-tuning. Specifically, we prioritize training on "mechanistically hard" samples-those that fail to naturally activate the translation initiation features. Experiments show this approach significantly improves data efficiency and suppresses hallucinations. Furthermore, we find these mechanisms are transferable to larger models of the same family. Our work not only decodes a core component of the translation mechanism in LLMs but also provides a blueprint for using internal model mechanism to create more robust and efficient models. The codes are available at https://github.com/flamewei123/AAAI26-translation- Initiation-Features.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c339f1f7-aa1c-4964-8fd2-5a80452bbd50Cited by top-tier papers2
- Incentivizing Parametric Knowledge via Reinforcement Learning with Verifiable Rewards for Cross-Cultural Entity TranslationJiang Zhou, Xiaohu Zhao, Xinwei Wu, Tianyu Dong et al.ACL 2026 · 2 citations
- From Insight to Action: A Novel Framework for Interpretability-Guided Data Selection in Large Language ModelsLing Shi, Xinwei Wu, Xiaohu Zhao, Hao Wang et al.ACL 2026
Builds on13
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Transcoders find interpretable LLM feature circuitsJacob Dunefsky, Philippe Chlenski, Neel NandaNeurIPS 2024 · 222 citations
- Task Contamination: Language Models May Not Be Few-Shot AnymoreChangmao Li, Jeffrey FlaniganAAAI 2024 · 138 citations
- Document-Level Machine Translation with Large Language ModelsLongyue Wang, Chenyang Lyu, Tianbo Ji, Zhirui Zhang et al.EMNLP 2023 · 129 citations
- Should We Really Edit Language Models? On the Evaluation of Edited Language ModelsQi Li, Xiang Liu, Zhenheng Tang, Peijie Dong et al.NeurIPS 2024 · 25 citations
Related papers
- Unveiling Language-Specific Features in Large Language Models via Sparse AutoencodersBoyi Deng, Yu Wan, Baosong Yang, Yidan Zhang et al.ACL 2025
- Sparse Autoencoders Trained on the Same Data Learn Different FeaturesGonçalo Paulo, Nora BelroseICLR 2026 · 96 citations
- Toward Faithful Retrieval-Augmented Generation with Sparse AutoencodersGuangzhi Xiong, Zhenghao He, Bohan Liu, Sanchit Sinha et al.ICLR 2026 · 8 citations
- Exploring the Translation Mechanism of Large Language ModelsHongbin Zhang, Kehai Chen, Xuefeng Bai, Xiucheng Li et al.NeurIPS 2025 · 4 citations
- DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse AutoencodersXu Wang, Bingqing Jiang, Yu Wan, Baosong Yang et al.ICML 2026
