Accurate and Robust Feature Importance Estimation under Distribution Shifts
Jayaraman J. Thiagarajan, Vivek Sivaraman Narayanaswamy, Rushil Anirudh, Peer-Timo Bremer, Andreas Spanias
Abstract
With increasing reliance on the outcomes of black-box models in critical applications, post-hoc explainability tools that do not require access to the model internals are often used to enable humans understand and trust these models. In particular, we focus on the class of methods that can reveal the influence of input features on the predicted outputs. Despite their wide-spread adoption, existing methods are known to suffer from one or more of the following challenges: computational complexities, large uncertainties and most importantly, inability to handle real-world domain shifts. In this paper, we propose PRoFILE (Producing Robust Feature Importances using Loss Estimates), a novel feature importance estimation method that addresses all these challenges. Through the use of a loss estimator jointly trained with the predictive model and a causal objective, PRoFILE can accurately estimate the feature importance scores even under complex distribution shifts, without any additional re-training. To this end, we also develop learning strategies for training the loss estimator, namely contrastive and dropout calibration, and find that it can effectively detect distribution shifts. Using empirical studies on several benchmark image and non-image data, we show significant improvements over state-of-the-art approaches, both in terms of fidelity and robustness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0858cdc8-5c15-4b1a-baee-b9fb5d69c3eeCited by top-tier papers1
Ask how each one uses itBuilds on2
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford et al.ICLR 2020 · 974 citations
- Building Calibrated Deep Models via Uncertainty Matching with Auxiliary Interval PredictorsJayaraman J. Thiagarajan, Bindya Venkatesh, Prasanna Sattigeri, Peer-Timo BremerAAAI 2020 · 35 citations
Related papers
- IW-GAE: Importance weighted group accuracy estimation for improved calibration and model selection in unsupervised domain adaptationTaejong Joo, Diego KlabjanICML 2024 · 1 citation
- Label-Free Explainability for Unsupervised ModelsJonathan Crabbé, Mihaela van der SchaarICML 2022 · 24 citations
- Explanatory Model Monitoring to Understand the Effects of Feature Shifts on PerformanceThomas Decker, Alexander Koebler, Michael Lebacher, Ingo Thon et al.KDD 2024
- Predicting with Confidence on Unseen DistributionsDevin Guillory, Vaishaal Shankar, Sayna Ebrahimi, Trevor Darrell et al.ICCV 2021 · 141 citations
- Robust and Stable Black Box ExplanationsHimabindu Lakkaraju, Nino Arsov, Osbert BastaniICML 2020 · 93 citations
