Addressing divergent representations from causal interventions on neural networks
Satchel Grant, Simon Jerome Han, Alexa R. Tartaglini, Christopher Potts
Abstract
A common approach to mechanistic interpretability is to causally manipulate model representations via targeted interventions in order to understand what those representations encode. Here we ask whether such interventions create out-ofdistribution (divergent) representations, and whether this raises concerns about how faithful their resulting explanations are to the target model in its natural state. First, we demonstrate theoretically and empirically that common causal intervention techniques often do shift internal representations away from the natural distribution of the target model. Then, we provide a theoretical analysis of two cases of such divergences: 'harmless' divergences that occur in the behavioral null-space of the layer(s) of interest, and 'pernicious' divergences that activate hidden network pathways and cause dormant behavioral changes. Finally, in an effort to mitigate the pernicious cases, we apply and modify the Counterfactual Latent (CL) loss from Grant (2025) allowing representations from causal interventions to remain closer to the natural distribution, reducing the likelihood of harmful divergences while preserving the interpretive power of the interventions. Together, these results highlight a path towards more reliable interpretability methods. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6bde6f53-7457-4cb6-aaa0-bd7faf3acd77Cited by top-tier papers1
Ask how each one uses itBuilds on13
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Causal Abstractions of Neural NetworksAtticus Geiger, Hanson Lu, Thomas Icard, Christopher PottsNeurIPS 2021 · 516 citations
- Towards Best Practices of Activation Patching in Language Models: Metrics and MethodsFred Zhang, Neel NandaICLR 2024 · 233 citations
- Interpretability at Scale: Identifying Causal Mechanisms in AlpacaZhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts et al.NeurIPS 2023 · 146 citations
- How do Language Models Bind Entities in Context?Jiahai Feng, Jacob SteinhardtICLR 2024 · 81 citations
Related papers
- Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation PatchingAleksandar Makelov, Georg Lange, Atticus Geiger, Neel NandaICLR 2024 · 48 citations
- Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution BehaviorsJing Huang, Junyi Tao, Thomas Icard, Diyi Yang et al.ICML 2025
- Latent Space Explanation by InterventionItai Gat, Guy Lorberbom, Idan Schwartz, Tamir HazanAAAI 2022 · 19 citations
- Counterfactual-based Saliency Map: Towards Visual Contrastive Explanations for Neural NetworksXue Wang, Zhibo Wang, Haiqin Weng, Hengchang Guo et al.ICCV 2023 · 15 citations
- Causal Differentiating Concepts: Interpreting LM Behavior via Causal Representation LearningNavita Goyal, Hal Daumé III, Alexandre Drouin, Dhanya SridharNeurIPS 2025 · 8 citations
