Successor Heads: Recurring, Interpretable Attention Heads In The Wild
Rhys Gould, Euan Ong, George Ogden, Arthur Conmy
Abstract
In this work we present successor heads: attention heads that increment tokens with a natural ordering, such as numbers, months, and days. For example, successor heads increment 'Monday' into 'Tuesday'. We explain the successor head behavior with an approach rooted in mechanistic interpretability, the field that aims to explain how models complete tasks in human-understandable terms. Existing research in this area has struggled to find recurring, mechanistically interpretable language model components beyond small toy models. Further, existing results have led to very little insight to explain the internals of larger models that are used in practice. In this paper, we analyze the behavior of successor heads in large language models (LLMs) and find that they implement abstract representations that are common to different architectures. They form in LLMs with as few as 31 million parameters, and at least as many as 12 billion parameters, such as GPT-2, Pythia, and Llama-2. We find a set of 'mod 10' features 1 that underlie how successor heads increment in LLMs across different architectures and sizes. We perform vector arithmetic with these features to edit head behavior and provide insights into numeric representations within LLMs. Finally, we study the behavior of successor heads on models' training data, finding that successor heads are important for the model getting low loss on examples of succession in this dataset. Finally, we interpret some of the other tasks these polysemantic heads perform and discuss the implications of our findings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e9cf0c8-1f27-49d1-8452-5ef310b11cd0Cited by top-tier papers33
- Transcoders find interpretable LLM feature circuitsJacob Dunefsky, Philippe Chlenski, Neel NandaNeurIPS 2024 · 222 citations
- LLM Circuit Analyses Are Consistent Across Training and ScaleCurt Tigges, Michael Hanna, Qinan Yu, Stella BidermanNeurIPS 2024 · 71 citations
- Do Language Models Use Their Depth Efficiently?Róbert Csordás, Christopher D. Manning, Christopher PottsNeurIPS 2025 · 61 citations
- Explorations of Self-Repair in Language ModelsCody Rushing, Neel NandaICML 2024 · 22 citations
- Feature Contamination: Neural Networks Learn Uncorrelated Features and Fail to GeneralizeTianren Zhang, Chujie Zhao, Guanyu Chen, Yizhou Jiang et al.ICML 2024 · 12 citations
Builds on5
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim et al.NeurIPS 2023 · 861 citations
- A Toy Model of Universality: Reverse Engineering how Networks Learn Group OperationsBilal Chughtai, Lawrence Chan, Neel NandaICML 2023 · 144 citations
- Progress measures for grokking via mechanistic interpretabilityNeel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith et al.ICLR 2023 · 54 citations
- Generalizing Backpropagation for Gradient-Based InterpretabilityKevin Du, Lucas Torroba Hennigen, Niklas Stoehr, Alex Warstadt et al.ACL 2023 · 3 citations
Related papers
- On the Role of Attention Heads in Large Language Model SafetyZhenhong Zhou, Haiyang Yu, Xinghua Zhang, Rongwu Xu et al.ICLR 2025
- Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in TransformersAndrew Nam, Henry Conklin, Yukang Yang, Tom Griffiths et al.NeurIPS 2025 · 21 citations
- Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language ModelsMichael Lan, Philip Torr, Fazl BarezEMNLP 2024 · 1 citation
- Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 SmallKevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris et al.ICLR 2023 · 50 citations
- The Same but Different: Structural Similarities and Differences in Multilingual Language ModelingRuochen Zhang, Qinan Yu, Matianyu Zang, Carsten Eickhoff et al.ICLR 2025
