Explorations of Self-Repair in Language Models
Cody Rushing, Neel Nanda
Abstract
Prior interpretability research studying narrow distributions has preliminarily identified self-repair, a phenomena where if components in large language models are ablated, later components will change their behavior to compensate. Our work builds off this past literature, demonstrating that self-repair exists on a variety of models families and sizes when ablating individual attention heads on the full training distribution. We further show that on the full training distribution self-repair is imperfect, as the original direct effect of the head is not fully restored, and noisy, since the degree of self-repair varies significantly across different prompts (sometimes overcorrecting beyond the original effect). We highlight two different mechanisms that contribute to self-repair, including changes in the final LayerNorm scaling factor and sparse sets of neurons implementing Anti-Erasure. We additionally discuss the implications of these results for interpretability practitioners and close with a more speculative discussion on the mystery of why self-repair occurs in these models at all, highlighting evidence for the Iterative Inference hypothesis in language models, a framework that predicts self-repair.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 14ee8e0b-abb2-4df9-abd9-aa2f79cbb68eCited by top-tier papers11
- LLM Circuit Analyses Are Consistent Across Training and ScaleCurt Tigges, Michael Hanna, Qinan Yu, Stella BidermanNeurIPS 2024 · 71 citations
- Optimal ablation for interpretabilityMaximilian Li, Lucas JansonNeurIPS 2024 · 32 citations
- Structural Inference: Interpreting Small Language Models with SusceptibilitiesGarrett Baker, George Wang, Jesse Hoogland, Vinayak Pathak et al.ICLR 2026 · 11 citations
- LLM Layers Immediately Correct Each OtherArjun Patrawala, Jiahai Feng, Erik Jones, Jacob SteinhardtNeurIPS 2025 · 5 citations
- Remarkable Robustness of LLMs: Stages of Inference?Vedang Lad, Jin Hwa Lee, Wes Gurnee, Max TegmarkNeurIPS 2025 · 5 citations
Builds on7
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim et al.NeurIPS 2023 · 861 citations
- Successor Heads: Recurring, Interpretable Attention Heads In The WildRhys Gould, Euan Ong, George Ogden, Arthur ConmyICLR 2024 · 75 citations
- Progress measures for grokking via mechanistic interpretabilityNeel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith et al.ICLR 2023 · 54 citations
Related papers
- Small Transformers Don’t Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and Implications for Mechanistic InterpretabilityLuca Baroni, Galvin Khara, Joachim Schaeffer, Marat Subkhankulov et al.ICLR 2026 · 8 citations
- Mechanisms of Prompt-Induced Hallucination in Vision-Language ModelsWilliam Rudman, Michal Golovanevsky, Dana Arad, Yonatan Belinkov et al.ACL 2026 · 4 citations
- Interpreting the Repeated Token Phenomenon in Large Language ModelsItay Yona, Ilia Shumailov, Jamie Hayes, Yossi GandelsmanICML 2025
- Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 SmallKevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris et al.ICLR 2023 · 50 citations
- Cognitive Mirrors: Exploring the Diverse Functional Roles of Attention Heads in LLM ReasoningXueqi Ma, Jun Wang, Yanbei Jiang, Sarah M. Erfani et al.NeurIPS 2025 · 5 citations
