Robust and Stable Black Box Explanations
Himabindu Lakkaraju, Nino Arsov, Osbert Bastani
Abstract
As machine learning black boxes are increasingly being deployed in real-world applications, there has been a growing interest in developing post hoc explanations that summarize the behaviors of these black boxes. However, existing algorithms for generating such explanations have been shown to lack stability and robustness to distribution shifts. We propose a novel framework for generating robust and stable explanations of black box models based on adversarial training. Our framework optimizes a minimax objective that aims to construct the highest fidelity explanation with respect to the worst-case over a set of adversarial perturbations. We instantiate this algorithm for explanations in the form of linear models and decision sets by devising the required optimization procedures. To the best of our knowledge, this work makes the first attempt at generating post hoc explanations that are robust to a general class of adversarial perturbations that are of practical interest. Experimental evaluation with real-world and synthetic datasets demonstrates that our approach substantially improves robustness of explanations without sacrificing their fidelity on the original data distribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 445ef5be-f5b6-46bb-b56b-da1ac6a4458dCited by top-tier papers20
- Towards Robust and Reliable Algorithmic RecourseSohini Upadhyay, Shalmali Joshi, Himabindu LakkarajuNeurIPS 2021 · 145 citations
- Post Hoc Explanations of Language Models Can Improve Language ModelsSatyapriya Krishna, Jiaqi Ma, Dylan Slack, Asma Ghandeharioun et al.NeurIPS 2023 · 87 citations
- Do Input Gradients Highlight Discriminative Features?Harshay Shah, Prateek Jain, Praneeth NetrapalliNeurIPS 2021 · 74 citations
- Provably efficient, succinct, and precise explanationsGuy Blanc, Jane Lange, Li-Yang TanNeurIPS 2021 · 45 citations
- A Framework for Learning Ante-hoc Explainable Models via ConceptsAnirban Sarkar, Deepak Vijaykeerthy, Anindya Sarkar, Vineeth N. BalasubramanianCVPR 2022 · 40 citations
Related papers
- SoK: Explainable Machine Learning in Adversarial EnvironmentsMaximilian Noppel, Christian WressneggerS&P 2024 · 28 citations
- Characterizing the risk of fairwashingUlrich Aïvodji, Hiromi Arai, Sébastien Gambs, Satoshi HaraNeurIPS 2021 · 35 citations
- Discriminative Feature Attributions: Bridging Post Hoc Explainability and Inherent InterpretabilityUsha Bhalla, Suraj Srinivas, Himabindu LakkarajuNeurIPS 2023 · 18 citations
- Towards the Unification and Robustness of Perturbation and Gradient Based ExplanationsSushant Agarwal, Shahin Jabbari, Chirag Agarwal, Sohini Upadhyay et al.ICML 2021 · 71 citations
- Robust Explanation Constraints for Neural NetworksMatthew Wicker, Juyeon Heo, Luca Costabello, Adrian WellerICLR 2023 · 3 citations
