Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
Fred Zhang, Neel Nanda
摘要
Mechanistic interpretability seeks to understand the internal mechanisms of machine learning models, where localization -- identifying the important model components -- is a key step. Activation patching, also known as causal tracing or interchange intervention, is a standard technique for this task (Vig et al., 2020), but the literature contains many variants with little consensus on the choice of hyperparameters or methodology. In this work, we systematically examine the impact of methodological details in activation patching, including evaluation metrics and corruption methods. In several settings of localization and circuit discovery in language models, we find that varying these hyperparameters could lead to disparate interpretability results. Backed by empirical observations, we give conceptual arguments for why certain metrics or methods may be preferred. Finally, we provide recommendations for the best practices of activation patching going forwards.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper88
- Transcoders find interpretable LLM feature circuitsJacob Dunefsky, Philippe Chlenski, Neel NandaNeurIPS 2024 · 被引用 222 次
- Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language ModelsAsma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon 等ICML 2024 · 被引用 197 次
- Finding Transformer Circuits With Edge PruningAdithya Bhaskar, Alexander Wettig, Dan Friedman, Danqi ChenNeurIPS 2024 · 被引用 72 次
- LLM Circuit Analyses Are Consistent Across Training and ScaleCurt Tigges, Michael Hanna, Qinan Yu, Stella BidermanNeurIPS 2024 · 被引用 71 次
- Language Models Are Capable of Metacognitive Monitoring and Control of Their Internal ActivationsJi-An Li, Huadong Xiong, Robert C. Wilson, Marcelo G. Mattar 等NeurIPS 2025 · 被引用 49 次
它引用的顶会 Paper27
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister 等NeurIPS 2023 · 被引用 1,549 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim 等NeurIPS 2023 · 被引用 861 次
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian 等NeurIPS 2020 · 被引用 851 次
相关 Paper
- MIB: A Mechanistic Interpretability BenchmarkAaron Mueller, Atticus Geiger, Sarah Wiegreffe, Dana Arad 等ICML 2025
- Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation PatchingAleksandar Makelov, Georg Lange, Atticus Geiger, Neel NandaICLR 2024 · 被引用 48 次
- Mechanistic Interpretability as Statistical Estimation: A Variance AnalysisMaxime Méloux, François Portet, Maxime PeyrardICML 2026 · 被引用 13 次
- EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit IdentificationLin Zhang, Wenshuo Dong, Zhuoran Zhang, Shu Yang 等NeurIPS 2025 · 被引用 26 次
- Towards Faithful Natural Language Explanations: A Study Using Activation Patching in Large Language ModelsWei Jie Yeo, Ranjan Satapathy, Erik CambriaEMNLP 2025 · 被引用 2 次
