Talking Heads: Understanding Inter-Layer Communication in Transformer Language Models
Jack Merullo, Carsten Eickhoff, Ellie Pavlick
Abstract
Although it is known that transformer language models (LMs) pass features from early layers to later layers, it is not well understood how this information is represented and routed by the model. We analyze a mechanism used in two LMs to selectively inhibit items in a context in one task, and find that it underlies a commonly used abstraction across many context-retrieval behaviors. Specifically, we find that models write into low-rank subspaces of the residual stream to represent features which are then read out by later layers, forming low-rank communication channels (Elhage et al., 2021) between layers. A particular 3D subspace in model activations in GPT-2 can be traversed to positionally index items in lists, and we show that this mechanism can explain an otherwise arbitrary-seeming sensitivity of the model to the order of items in the prompt. That is, the model has trouble copying the correct information from context when many items ``crowd"this limited space. By decomposing attention heads with the Singular Value Decomposition (SVD), we find that previously described interactions between heads separated by one or more layers can be predicted via analysis of their weight matrices alone. We show that it is possible to manipulate the internal model representations as well as edit model weights based on the mechanism we discover in order to significantly improve performance on our synthetic Laundry List task, which requires recall from a list, often improving task accuracy by over 20%. Our analysis reveals a surprisingly intricate interpretable structure learned from language model pretraining, and helps us understand why sophisticated LMs sometimes fail in simple domains, facilitating future analysis of more complex behaviors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 272f0675-bb1b-4d3d-ac3c-ce2970c102d3Cited by top-tier papers23
- Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in TransformersAndrew Nam, Henry Conklin, Yukang Yang, Tom Griffiths et al.NeurIPS 2025 · 21 citations
- Head Pursuit: Probing Attention Specialization in Multimodal TransformersLorenzo Basile, Valentino Maiorca, Diego Doimo, Francesco Locatello et al.NeurIPS 2025 · 21 citations
- Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target AtomsMengru Wang, Ziwen Xu, Shengyu Mao, Shumin Deng et al.ACL 2025 · 19 citations
- A Implies B: Circuit Analysis in LLMs for Propositional Logical ReasoningGuanzhe Hong, Nishanth Dikkala, Enming Luo, Cyrus Rashtchian et al.NeurIPS 2025 · 17 citations
- Beyond Components: Singular Vector-Based Interpretability of Transformer CircuitsAreeb Ahmad, Abhinav Joshi, Ashutosh ModiNeurIPS 2025 · 9 citations
Builds on23
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel et al.ACL 2022 · 1,494 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Compacter: Efficient Low-Rank Hypercomplex Adapter LayersRabeeh Karimi Mahabadi, James Henderson, Sebastian RuderNeurIPS 2021 · 700 citations
Related papers
- A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation AnalysisAlessandro Stolfo, Yonatan Belinkov, Mrinmaya SachanEMNLP 2023 · 11 citations
- Pinpointing Attention-Causal Communication in Language ModelsGabriel Franco, Mark CrovellaNeurIPS 2025 · 3 citations
- Residual Stream Analysis with Multi-Layer SAEsTim Lawson, Lucy Farnik, Conor J. Houghton, Laurence AitchisonICLR 2025
- LLM Layers Immediately Correct Each OtherArjun Patrawala, Jiahai Feng, Erik Jones, Jacob SteinhardtNeurIPS 2025 · 5 citations
- LLMs Process Lists With General Filter HeadsArnab Sen Sharma, Giordano Rogers, Natalie Shapira, David BauICLR 2026 · 7 citations
