Successor Heads: Recurring, Interpretable Attention Heads In The Wild
Rhys Gould, Euan Ong, George Ogden, Arthur Conmy
摘要
In this work we present successor heads: attention heads that increment tokens with a natural ordering, such as numbers, months, and days. For example, successor heads increment 'Monday' into 'Tuesday'. We explain the successor head behavior with an approach rooted in mechanistic interpretability, the field that aims to explain how models complete tasks in human-understandable terms. Existing research in this area has struggled to find recurring, mechanistically interpretable language model components beyond small toy models. Further, existing results have led to very little insight to explain the internals of larger models that are used in practice. In this paper, we analyze the behavior of successor heads in large language models (LLMs) and find that they implement abstract representations that are common to different architectures. They form in LLMs with as few as 31 million parameters, and at least as many as 12 billion parameters, such as GPT-2, Pythia, and Llama-2. We find a set of 'mod 10' features 1 that underlie how successor heads increment in LLMs across different architectures and sizes. We perform vector arithmetic with these features to edit head behavior and provide insights into numeric representations within LLMs. Finally, we study the behavior of successor heads on models' training data, finding that successor heads are important for the model getting low loss on examples of succession in this dataset. Finally, we interpret some of the other tasks these polysemantic heads perform and discuss the implications of our findings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper33
- Transcoders find interpretable LLM feature circuitsJacob Dunefsky, Philippe Chlenski, Neel NandaNeurIPS 2024 · 被引用 222 次
- LLM Circuit Analyses Are Consistent Across Training and ScaleCurt Tigges, Michael Hanna, Qinan Yu, Stella BidermanNeurIPS 2024 · 被引用 71 次
- Do Language Models Use Their Depth Efficiently?Róbert Csordás, Christopher D. Manning, Christopher PottsNeurIPS 2025 · 被引用 61 次
- Explorations of Self-Repair in Language ModelsCody Rushing, Neel NandaICML 2024 · 被引用 22 次
- Feature Contamination: Neural Networks Learn Uncorrelated Features and Fail to GeneralizeTianren Zhang, Chujie Zhao, Guanyu Chen, Yizhou Jiang 等ICML 2024 · 被引用 12 次
它引用的顶会 Paper5
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim 等NeurIPS 2023 · 被引用 861 次
- A Toy Model of Universality: Reverse Engineering how Networks Learn Group OperationsBilal Chughtai, Lawrence Chan, Neel NandaICML 2023 · 被引用 144 次
- Progress measures for grokking via mechanistic interpretabilityNeel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith 等ICLR 2023 · 被引用 54 次
- Generalizing Backpropagation for Gradient-Based InterpretabilityKevin Du, Lucas Torroba Hennigen, Niklas Stoehr, Alex Warstadt 等ACL 2023 · 被引用 3 次
相关 Paper
- On the Role of Attention Heads in Large Language Model SafetyZhenhong Zhou, Haiyang Yu, Xinghua Zhang, Rongwu Xu 等ICLR 2025
- Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in TransformersAndrew Nam, Henry Conklin, Yukang Yang, Tom Griffiths 等NeurIPS 2025 · 被引用 21 次
- Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language ModelsMichael Lan, Philip Torr, Fazl BarezEMNLP 2024 · 被引用 1 次
- Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 SmallKevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris 等ICLR 2023 · 被引用 50 次
- The Same but Different: Structural Similarities and Differences in Multilingual Language ModelingRuochen Zhang, Qinan Yu, Matianyu Zang, Carsten Eickhoff 等ICLR 2025
