Inducing Causal Structure for Interpretable Neural Networks
Atticus Geiger, Zhengxuan Wu, Hanson Lu, Josh Rozner, Elisa Kreiss, Thomas Icard, Noah D. Goodman, Christopher Potts
Abstract
In many areas, we have well-founded insights about causal structure that would be useful to bring into our trained models while still allowing them to learn in a data-driven fashion. To achieve this, we present the new method of interchange intervention training (IIT). In IIT, we (1) align variables in a causal model (e.g., a deterministic program or Bayesian network) with representations in a neural model and ( 2 ) train the neural model to match the counterfactual behavior of the causal model on a base input when aligned representations in both models are set to be the value they would be for a source input. IIT is fully differentiable, flexibly combines with other objectives, and guarantees that the target causal model is a causal abstraction of the neural model when its loss is zero. We evaluate IIT on a structural vision task (MNIST-PVR), a navigational language task (ReaSCAN), and a natural language inference task (MQNLI). We compare IIT against multi-task training objectives and data augmentation. In all our experiments, IIT achieves the best results and produces neural models that are more interpretable in the sense that they more successfully realize the target causal model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers28
- LEACE: Perfect linear concept erasure in closed formNora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell et al.NeurIPS 2023 · 305 citations
- ReFT: Representation Finetuning for Language ModelsZhengxuan Wu, Aryaman Arora, Zheng Wang, Atticus Geiger et al.NeurIPS 2024 · 233 citations
- Towards Best Practices of Activation Patching in Language Models: Metrics and MethodsFred Zhang, Neel NandaICLR 2024 · 233 citations
- Interpretability at Scale: Identifying Causal Mechanisms in AlpacaZhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts et al.NeurIPS 2023 · 146 citations
- Language Models Use Lookbacks to Track BeliefsNikhil Prakash, Natalie Shapira, Arnab Sen Sharma, Christoph Riedl et al.ICLR 2026 · 42 citations
Builds on7
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 625 citations
- Causal Abstractions of Neural NetworksAtticus Geiger, Hanson Lu, Thomas Icard, Christopher PottsNeurIPS 2021 · 516 citations
- A Benchmark for Systematic Generalization in Grounded Language UnderstandingLaura Ruis, Jacob Andreas, Marco Baroni, Diane Bouchacourt et al.NeurIPS 2020 · 169 citations
- Counterfactuals uncover the modular structure of deep generative modelsMichel Besserve, Arash Mehrjou, Rémy Sun, Bernhard SchölkopfICLR 2020 · 109 citations
Related papers
- Counterfactual Vision and Language LearningEhsan Abbasnejad, Damien Teney, Amin Parvaneh, Javen Shi et al.CVPR 2020
- From Prompts to Tokens: Internalizing Causal Supervision in Vision-Language Model for Multi-Image Causal ReasoningHaoping Yu, Yuanxi Li, Jing MaICML 2026
- The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?Denis Sutter, Julian Minder, Thomas Hofmann, Tiago PimentelNeurIPS 2025 · 30 citations
- Neural Causal Graph for Interpretable and Intervenable ClassificationJiawei Wang, Shaofei Lu, Da Cao, Dongyu Wang et al.ICLR 2025
- Generative Interventions for Causal LearningChengzhi Mao, Augustine Cha, Amogh Gupta, Hao Wang et al.CVPR 2021
