The Hidden Space of Transformer Language Adapters
Jesujoba Alabi, Marius Mosbach, Matan Eyal, Dietrich Klakow, Mor Geva
摘要
We analyze the operation of transformer language adapters, which are small modules trained on top of a frozen language model to adapt its predictions to new target languages. We show that adapted predictions mostly evolve in the source language the model was trained on, while the target language becomes pronounced only in the very last layers of the model. Moreover, the adaptation process is gradual and distributed across layers, where it is possible to skip small groups of adapters without decreasing adaptation performance. Last, we show that adapters operate on top of the model's frozen representation space while largely preserving its structure, rather than on an "isolated" subspace. Our findings provide a deeper view into the adaptation process of language models to new languages, showcasing the constraints imposed on it by the underlying model and introduces practical implications to enhance its efficiency. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Multilingual Routing in Mixture-of-ExpertsLucas Bandarkar, Chenyuan Yang, Mohsen Fayyaz, Junlin Hu 等ICLR 2026 · 被引用 34 次
- Beyond the Surface: Enhancing LLM-as-a-Judge Alignment with Human via Internal RepresentationsPeng Lai, Jianjie Zheng, Sijie Cheng, Yun Chen 等NeurIPS 2025 · 被引用 16 次
- Layer Swapping for Zero-Shot Cross-Lingual Transfer in Large Language ModelsLucas Bandarkar, Benjamin Muller, Pritish Yuvraj, Rui Hou 等ICLR 2025
它引用的顶会 Paper15
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Neural Machine Translation with Byte-Level SubwordsChanghan Wang, Kyunghyun Cho, Jiatao GuAAAI 2020 · 被引用 213 次
- Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary SpaceMor Geva, Avi Caciularu, Kevin Ro Wang, Yoav GoldbergEMNLP 2022 · 被引用 92 次
- On the Cross-lingual Transferability of Monolingual RepresentationsMikel Artetxe, Sebastian Ruder, Dani YogatamaACL 2020 · 被引用 57 次
相关 Paper
- Transformer Layers as PaintersQi Sun, Marc Pickett, Aakash Kumar Nain, Llion JonesAAAI 2025 · 被引用 49 次
- Frozen Pretrained Transformers as Universal Computation EnginesKevin Lu, Aditya Grover, Pieter Abbeel, Igor MordatchAAAI 2022 · 被引用 133 次
- IAA: Inner-Adaptor Architecture Empowers Frozen Large Language Model with Multimodal CapabilitiesBin Wang, Chunyu Xie, Dawei Leng, Yuhui YinAAAI 2025 · 被引用 8 次
- Attentive Multi-Layer Fusion for Vision TransformersLaure Ciernik, Marco Morik, Lukas Thede, Luca Eyring 等ICML 2026 · 被引用 3 次
- Separating Tongue from Thought: Activation Patching Reveals Language-Agnostic Concept Representations in TransformersClément Dumas, Chris Wendler, Veniamin Veselovsky, Giovanni Monea 等ACL 2025
