Right for Better Reasons: Training Differentiable Models by Constraining their Influence Functions
Xiaoting Shao, Arseny Skryagin, Wolfgang Stammer, Patrick Schramowski, Kristian Kersting
摘要
Explaining black-box models such as deep neural networks is becoming increasingly important as it helps to boost trust and debugging. Popular forms of explanations map the features to a vector indicating their individual importance to a decision on the instance-level. They can then be used to prevent the model from learning the wrong bias in data possibly due to ambiguity. For instance, Ross et al.'s "right for the right reasons" propagates user explanations backwards to the network by formulating differentiable constraints based on input gradients. Unfortunately, input gradients as well as many other widely used explanation methods form an approximation of the decision boundary and assume the underlying model to be fixed. Here, we demonstrate how to make use of influence functions-a well known robust statistic-in the constraints to correct the model's behaviour more effectively. Our empirical evidence demonstrates that this "right for better reasons"(RBR) considerably reduces the time to correct the classifier at training time and boosts the quality of explanations at inference time compared to input gradients. Besides, we also showcase the effectiveness of RBR in correcting "Clever Hans"-like behaviour in real, high-dimensional domain.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- A Rationale-Centric Framework for Human-in-the-loop Machine LearningJinghui Lu, Linyi Yang, Brian MacNamee, Yue ZhangACL 2022 · 被引用 46 次
- Studying How to Efficiently and Effectively Guide Models with ExplanationsSukrut Rao, Moritz Böhle, Amin Parchami-Araghi, Bernt SchieleICCV 2023 · 被引用 22 次
- Targeted Activation Penalties Help CNNs Ignore Spurious SignalsDekai Zhang, Matt Williams, Francesca ToniAAAI 2024 · 被引用 3 次
- Use perturbations when learning from explanationsJuyeon Heo, Vihari Piratla, Matthew Wicker, Adrian WellerNeurIPS 2023 · 被引用 2 次
- Meta-Sift: How to Sift Out a Clean Subset in the Presence of Data Poisoning?Yi Zeng, Minzhou Pan, Himanshu Jahagirdar, Ming Jin 等USENIX Security 2023
它引用的顶会 Paper1
相关 Paper
- Robust Explanation Constraints for Neural NetworksMatthew Wicker, Juyeon Heo, Luca Costabello, Adrian WellerICLR 2023 · 被引用 3 次
- Learning with Explanation ConstraintsRattana Pukdee, Dylan Sam, J. Zico Kolter, Maria-Florina Balcan 等NeurIPS 2023 · 被引用 11 次
- Explaining Neural Matrix Factorization with Gradient RollbackCarolin Lawrence, Timo Sztyler, Mathias NiepertAAAI 2021 · 被引用 15 次
- Theoretical and Practical Perspectives on what Influence Functions DoAndrea Schioppa, Katja Filippova, Ivan Titov, Polina ZablotskaiaNeurIPS 2023 · 被引用 38 次
- Learning Deep Attribution Priors Based On Prior KnowledgeEthan Weinberger, Joseph D. Janizek, Su-In LeeNeurIPS 2020 · 被引用 27 次
