Explorations of Self-Repair in Language Models
Cody Rushing, Neel Nanda
摘要
Prior interpretability research studying narrow distributions has preliminarily identified self-repair, a phenomena where if components in large language models are ablated, later components will change their behavior to compensate. Our work builds off this past literature, demonstrating that self-repair exists on a variety of models families and sizes when ablating individual attention heads on the full training distribution. We further show that on the full training distribution self-repair is imperfect, as the original direct effect of the head is not fully restored, and noisy, since the degree of self-repair varies significantly across different prompts (sometimes overcorrecting beyond the original effect). We highlight two different mechanisms that contribute to self-repair, including changes in the final LayerNorm scaling factor and sparse sets of neurons implementing Anti-Erasure. We additionally discuss the implications of these results for interpretability practitioners and close with a more speculative discussion on the mystery of why self-repair occurs in these models at all, highlighting evidence for the Iterative Inference hypothesis in language models, a framework that predicts self-repair.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- LLM Circuit Analyses Are Consistent Across Training and ScaleCurt Tigges, Michael Hanna, Qinan Yu, Stella BidermanNeurIPS 2024 · 被引用 71 次
- Optimal ablation for interpretabilityMaximilian Li, Lucas JansonNeurIPS 2024 · 被引用 32 次
- Structural Inference: Interpreting Small Language Models with SusceptibilitiesGarrett Baker, George Wang, Jesse Hoogland, Vinayak Pathak 等ICLR 2026 · 被引用 11 次
- LLM Layers Immediately Correct Each OtherArjun Patrawala, Jiahai Feng, Erik Jones, Jacob SteinhardtNeurIPS 2025 · 被引用 5 次
- Remarkable Robustness of LLMs: Stages of Inference?Vedang Lad, Jin Hwa Lee, Wes Gurnee, Max TegmarkNeurIPS 2025 · 被引用 5 次
它引用的顶会 Paper7
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim 等NeurIPS 2023 · 被引用 861 次
- Successor Heads: Recurring, Interpretable Attention Heads In The WildRhys Gould, Euan Ong, George Ogden, Arthur ConmyICLR 2024 · 被引用 75 次
- Progress measures for grokking via mechanistic interpretabilityNeel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith 等ICLR 2023 · 被引用 54 次
相关 Paper
- Small Transformers Don’t Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and Implications for Mechanistic InterpretabilityLuca Baroni, Galvin Khara, Joachim Schaeffer, Marat Subkhankulov 等ICLR 2026 · 被引用 8 次
- Mechanisms of Prompt-Induced Hallucination in Vision-Language ModelsWilliam Rudman, Michal Golovanevsky, Dana Arad, Yonatan Belinkov 等ACL 2026 · 被引用 4 次
- Interpreting the Repeated Token Phenomenon in Large Language ModelsItay Yona, Ilia Shumailov, Jamie Hayes, Yossi GandelsmanICML 2025
- Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 SmallKevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris 等ICLR 2023 · 被引用 50 次
- Cognitive Mirrors: Exploring the Diverse Functional Roles of Attention Heads in LLM ReasoningXueqi Ma, Jun Wang, Yanbei Jiang, Sarah M. Erfani 等NeurIPS 2025 · 被引用 5 次
