Positional Cognitive Specialization: Where Do LLMs Learn to Comprehend and Speak Your Language?
Luis Frentzen Salim, Lun-Wei Ku, Hsing-Kuo Kenneth Pao
摘要
Adapting large language models (LLMs) to new languages is an expensive and opaque process. Understanding how language models acquire new languages and multilingual abilities is key to achieve efficient adaptation. Prior work on multilingual interpretability research focuses primarily on how trained models process multilingual instructions, leaving unexplored the mechanisms through which they acquire new languages during training. We investigate these training dynamics on decoder-only transformers through the lens of two functional cognitive specializations: language perception (input comprehension) and production (output generation). Through experiments on low-resource languages, we demonstrate how perceptual and productive specialization emerges in different regions of a language model by running layer ablation sweeps from the model’s input and output directions. Based on the observed specialization patterns, we propose CogSym, a layer-wise heuristic that enables effective adaptation by exclusively finetuning a few early and late layers. We show that tuning only the 25% outermost layers achieves downstream task performance within 2–3% deviation from the full finetuning baseline. Unlike similar layer-selection methods, the proposed method requires no extra data or computation while retaining comparable performance, which is especially beneficial for low-resource languages. CogSym yields consistent performance with adapter methods such as LoRA, showcasing generalization beyond full finetuning. These findings provide insights to better understand how LLMs learn new languages and push toward accessible and inclusive language modeling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- How do Large Language Models Handle Multilingualism?Yiran Zhao, Wenxuan Zhang, Guizhen Chen, Kenji Kawaguchi 等NeurIPS 2024 · 被引用 196 次
- Few-shot Learning with Multilingual Generative Language ModelsXi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang 等EMNLP 2022 · 被引用 113 次
- Language models are multilingual chain-of-thought reasonersFreda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang 等ICLR 2023 · 被引用 52 次
相关 Paper
- Vision Function Layer in Multimodal LLMsCheng Shi, Yizhou Yu, Sibei YangNeurIPS 2025 · 被引用 20 次
- LinguaMap: Which Layers of LLMs Speak Your Language and How to Tune Them?J. Ben Tamo, Daniel Carlander-Reuterfelt, Jonathan Rubin, Oleg Poliannikov 等ICLR 2026 · 被引用 3 次
- MLAS-LoRA: Language-Aware Parameters Detection and LoRA-Based Knowledge Transfer for Multilingual Machine TranslationTianyu Dong, Bo Li, Jinsong Liu, Shaolin Zhu 等ACL 2025
- When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning MethodBiao Zhang, Zhongtao Liu, Colin Cherry, Orhan FiratICLR 2024 · 被引用 271 次
- Exploring Intrinsic Language-specific Subspaces in Fine-tuning Multilingual Neural Machine TranslationZhe Cao, Zhi Qu, Hidetaka Kamigaito, Taro WatanabeEMNLP 2024 · 被引用 1 次
