Post hoc Explanations may be Ineffective for Detecting Unknown Spurious Correlation
Julius Adebayo, Michael Muelly, Harold Abelson, Been Kim
Abstract
We investigate whether three types of post hoc model explanations-feature attribution, concept activation, and training point ranking-are effective for detecting a model's reliance on spurious signals in the training data. Specifically, we consider the scenario where the spurious signal to be detected is unknown, at test-time, to the user of the explanation method. We design an empirical methodology that uses semi-synthetic datasets along with pre-specified spurious artifacts to obtain models that verifiably rely on these spurious training signals. We then provide a suite of metrics that assess an explanation method's reliability for spurious signal detection under various conditions. We find that the post hoc explanation methods tested are ineffective when the spurious artifact is unknown at test-time especially for non-visible artifacts like a background blur. Further, we find that feature attribution methods are susceptible to erroneously indicating dependence on spurious signals even when the model being explained does not rely on spurious artifacts. This finding casts doubt on the utility of these approaches, in the hands of a practitioner, for detecting a model's reliance on spurious signals. 1 It is hard to find a needle in a haystack, it is much harder if you haven't seen a needle before (Pearl). -Judea Pearl
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- Spurious Features Everywhere - Large-Scale Detection of Harmful Spurious Features in ImageNetYannic Neuhaus, Maximilian Augustin, Valentyn Boreiko, Matthias HeinICCV 2023 · 42 citations
- "Will You Find These Shortcuts?" A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text ClassificationJasmijn Bastings, Sebastian Ebert, Polina Zablotskaia, Anders Sandholm et al.EMNLP 2022 · 29 citations
- Studying How to Efficiently and Effectively Guide Models with ExplanationsSukrut Rao, Moritz Böhle, Amin Parchami-Araghi, Bernt SchieleICCV 2023 · 22 citations
- On the Relationship Between Explanation and Prediction: A Causal ViewAmir-Hossein Karimi, Krikamol Muandet, Simon Kornblith, Bernhard Schölkopf et al.ICML 2023 · 20 citations
- Stability Guarantees for Feature Attributions with Multiplicative SmoothingAnton Xue, Rajeev Alur, Eric WongNeurIPS 2023 · 18 citations
Builds on18
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- Manipulating and Measuring Model InterpretabilityForough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan et al.CHI 2021 · 663 citations
- Noise or Signal: The Role of Image Backgrounds in Object RecognitionKai Yuanqing Xiao, Logan Engstrom, Andrew Ilyas, Aleksander MadryICLR 2021 · 451 citations
- An Investigation of Why Overparameterization Exacerbates Spurious CorrelationsShiori Sagawa, Aditi Raghunathan, Pang Wei Koh, Percy LiangICML 2020 · 436 citations
- On Completeness-aware Concept-Based Explanations in Deep Neural NetworksChih-Kuan Yeh, Been Kim, Sercan Ömer Arik, Chun-Liang Li et al.NeurIPS 2020 · 390 citations
Related papers
- Debugging Tests for Model ExplanationsJulius Adebayo, Michael Muelly, Ilaria Liccardi, Been KimNeurIPS 2020 · 209 citations
- An Empirical Study on Explanations in Out-of-Domain SettingsGeorge Chrysostomou, Nikolaos AletrasACL 2022
- Rethinking Explanation Evaluation Under the Retraining SchemeYi Cai, Thibaud Ardoin, Mayank Gulati, Gerhard WunderAAAI 2026
- Are All Spurious Features in Natural Language Alike? An Analysis through a Causal LensNitish Joshi, Xiang Pan, He HeEMNLP 2022 · 19 citations
- Spuriosity Rankings: Sorting Data to Measure and Mitigate BiasesMazda Moayeri, Wenxiao Wang, Sahil Singla, Soheil FeiziNeurIPS 2023 · 19 citations
