Addressing divergent representations from causal interventions on neural networks
Satchel Grant, Simon Jerome Han, Alexa R. Tartaglini, Christopher Potts
摘要
A common approach to mechanistic interpretability is to causally manipulate model representations via targeted interventions in order to understand what those representations encode. Here we ask whether such interventions create out-ofdistribution (divergent) representations, and whether this raises concerns about how faithful their resulting explanations are to the target model in its natural state. First, we demonstrate theoretically and empirically that common causal intervention techniques often do shift internal representations away from the natural distribution of the target model. Then, we provide a theoretical analysis of two cases of such divergences: 'harmless' divergences that occur in the behavioral null-space of the layer(s) of interest, and 'pernicious' divergences that activate hidden network pathways and cause dormant behavioral changes. Finally, in an effort to mitigate the pernicious cases, we apply and modify the Counterfactual Latent (CL) loss from Grant (2025) allowing representations from causal interventions to remain closer to the natural distribution, reducing the likelihood of harmful divergences while preserving the interpretive power of the interventions. Together, these results highlight a path towards more reliable interpretability methods. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper13
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Causal Abstractions of Neural NetworksAtticus Geiger, Hanson Lu, Thomas Icard, Christopher PottsNeurIPS 2021 · 被引用 516 次
- Towards Best Practices of Activation Patching in Language Models: Metrics and MethodsFred Zhang, Neel NandaICLR 2024 · 被引用 233 次
- Interpretability at Scale: Identifying Causal Mechanisms in AlpacaZhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts 等NeurIPS 2023 · 被引用 146 次
- How do Language Models Bind Entities in Context?Jiahai Feng, Jacob SteinhardtICLR 2024 · 被引用 81 次
相关 Paper
- Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation PatchingAleksandar Makelov, Georg Lange, Atticus Geiger, Neel NandaICLR 2024 · 被引用 48 次
- Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution BehaviorsJing Huang, Junyi Tao, Thomas Icard, Diyi Yang 等ICML 2025
- Latent Space Explanation by InterventionItai Gat, Guy Lorberbom, Idan Schwartz, Tamir HazanAAAI 2022 · 被引用 19 次
- Counterfactual-based Saliency Map: Towards Visual Contrastive Explanations for Neural NetworksXue Wang, Zhibo Wang, Haiqin Weng, Hengchang Guo 等ICCV 2023 · 被引用 15 次
- Causal Differentiating Concepts: Interpreting LM Behavior via Causal Representation LearningNavita Goyal, Hal Daumé III, Alexandre Drouin, Dhanya SridharNeurIPS 2025 · 被引用 8 次
