Right for Better Reasons: Training Differentiable Models by Constraining their Influence Functions
Xiaoting Shao, Arseny Skryagin, Wolfgang Stammer, Patrick Schramowski, Kristian Kersting
Abstract
Explaining black-box models such as deep neural networks is becoming increasingly important as it helps to boost trust and debugging. Popular forms of explanations map the features to a vector indicating their individual importance to a decision on the instance-level. They can then be used to prevent the model from learning the wrong bias in data possibly due to ambiguity. For instance, Ross et al.'s "right for the right reasons" propagates user explanations backwards to the network by formulating differentiable constraints based on input gradients. Unfortunately, input gradients as well as many other widely used explanation methods form an approximation of the decision boundary and assume the underlying model to be fixed. Here, we demonstrate how to make use of influence functions-a well known robust statistic-in the constraints to correct the model's behaviour more effectively. Our empirical evidence demonstrates that this "right for better reasons"(RBR) considerably reduces the time to correct the classifier at training time and boosts the quality of explanations at inference time compared to input gradients. Besides, we also showcase the effectiveness of RBR in correcting "Clever Hans"-like behaviour in real, high-dimensional domain.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 32a0f334-db4f-4ea1-86fb-7a8b537a1521Cited by top-tier papers5
- A Rationale-Centric Framework for Human-in-the-loop Machine LearningJinghui Lu, Linyi Yang, Brian MacNamee, Yue ZhangACL 2022 · 46 citations
- Studying How to Efficiently and Effectively Guide Models with ExplanationsSukrut Rao, Moritz Böhle, Amin Parchami-Araghi, Bernt SchieleICCV 2023 · 22 citations
- Targeted Activation Penalties Help CNNs Ignore Spurious SignalsDekai Zhang, Matt Williams, Francesca ToniAAAI 2024 · 3 citations
- Use perturbations when learning from explanationsJuyeon Heo, Vihari Piratla, Matthew Wicker, Adrian WellerNeurIPS 2023 · 2 citations
- Meta-Sift: How to Sift Out a Clean Subset in the Presence of Data Poisoning?Yi Zeng, Minzhou Pan, Himanshu Jahagirdar, Ming Jin et al.USENIX Security 2023
Builds on1
Related papers
- Robust Explanation Constraints for Neural NetworksMatthew Wicker, Juyeon Heo, Luca Costabello, Adrian WellerICLR 2023 · 3 citations
- Learning with Explanation ConstraintsRattana Pukdee, Dylan Sam, J. Zico Kolter, Maria-Florina Balcan et al.NeurIPS 2023 · 11 citations
- Explaining Neural Matrix Factorization with Gradient RollbackCarolin Lawrence, Timo Sztyler, Mathias NiepertAAAI 2021 · 15 citations
- Theoretical and Practical Perspectives on what Influence Functions DoAndrea Schioppa, Katja Filippova, Ivan Titov, Polina ZablotskaiaNeurIPS 2023 · 38 citations
- Learning Deep Attribution Priors Based On Prior KnowledgeEthan Weinberger, Joseph D. Janizek, Su-In LeeNeurIPS 2020 · 27 citations
