Matching Learned Causal Effects of Neural Networks with Domain Priors
Sai Srinivas Kancheti, Abbavaram Gowtham Reddy, Vineeth N. Balasubramanian, Amit Sharma
Abstract
A trained neural network can be interpreted as a structural causal model (SCM) that provides the effect of changing input variables on the model’s output. However, if training data contains both causal and correlational relationships, a model that optimizes prediction accuracy may not neces-sarily learn the true causal relationships between input and output variables. On the other hand, expert users often have prior knowledge of the causal relationship between certain input variables and output from domain knowledge. Therefore, we propose a regularization method that aligns the learned causal effects of a neural network with domain priors, including both direct and total causal effects. We show that this approach can generalize to different kinds of domain priors, including monotonicity of causal effect of an input variable on output or zero causal effect of a variable on output for purposes of fairness. Our experiments on twelve benchmark datasets show its utility in regularizing a neural network model to maintain desired causal effects, without compromising on accuracy. Importantly, we also show that a model thus trained is robust and gets improved accuracy on noisy inputs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext efdc2043-e470-4102-8767-1cc9b8b6ef78Cited by top-tier papers2
- Probing Classifiers are Unreliable for Concept Removal and DetectionAbhinav Kumar, Chenhao Tan, Amit SharmaNeurIPS 2022 · 46 citations
- Towards Learning and Explaining Indirect Causal Effects in Neural NetworksAbbavaram Gowtham Reddy, Saketh Bachu, Harsharaj Pathak, Benin Godfrey L et al.AAAI 2024 · 3 citations
Builds on6
- Causal Discovery with Reinforcement LearningShengyu Zhu, Ignavier Ng, Zhitang ChenICLR 2020 · 285 citations
- Interpretations are Useful: Penalizing Explanations to Align Neural Networks with Prior KnowledgeLaura Rieger, Chandan Singh, W. James Murdoch, Bin YuICML 2020 · 249 citations
- Learning Counterfactual Representations for Estimating Individual Dose-Response CurvesPatrick Schwab, Lorenz Linhardt, Stefan Bauer, Joachim M. Buhmann et al.AAAI 2020 · 159 citations
- Counterfactual Data Augmentation using Locally Factored DynamicsSilviu Pitis, Elliot Creager, Animesh GargNeurIPS 2020 · 126 citations
- CASTLE: Regularization via Auxiliary Causal Graph DiscoveryTrent Kyono, Yao Zhang, Mihaela van der SchaarNeurIPS 2020 · 82 citations
Related papers
- Counterfactual Maximum Likelihood Estimation for Training Deep NetworksXinyi Wang, Wenhu Chen, Michael Saxon, William Yang WangNeurIPS 2021 · 9 citations
- Inducing Causal Structure for Interpretable Neural NetworksAtticus Geiger, Zhengxuan Wu, Hanson Lu, Josh Rozner et al.ICML 2022 · 104 citations
- Foundation Models for Causal Inference via Prior-Data Fitted NetworksYuchen Ma, Dennis Frauen, Emil Javurek, Stefan FeuerriegelICLR 2026 · 37 citations
- From Predictions to Decisions: Using Lookahead RegularizationNir Rosenfeld, Sophie Hilgard, Sai Srivatsa Ravindranath, David C. ParkesNeurIPS 2020 · 27 citations
- Incorporating Interpretable Output Constraints in Bayesian Neural NetworksWanqian Yang, Lars Lorch, Moritz A. Graule, Himabindu Lakkaraju et al.NeurIPS 2020 · 17 citations
