Post hoc Explanations may be Ineffective for Detecting Unknown Spurious Correlation
Julius Adebayo, Michael Muelly, Harold Abelson, Been Kim
摘要
We investigate whether three types of post hoc model explanations-feature attribution, concept activation, and training point ranking-are effective for detecting a model's reliance on spurious signals in the training data. Specifically, we consider the scenario where the spurious signal to be detected is unknown, at test-time, to the user of the explanation method. We design an empirical methodology that uses semi-synthetic datasets along with pre-specified spurious artifacts to obtain models that verifiably rely on these spurious training signals. We then provide a suite of metrics that assess an explanation method's reliability for spurious signal detection under various conditions. We find that the post hoc explanation methods tested are ineffective when the spurious artifact is unknown at test-time especially for non-visible artifacts like a background blur. Further, we find that feature attribution methods are susceptible to erroneously indicating dependence on spurious signals even when the model being explained does not rely on spurious artifacts. This finding casts doubt on the utility of these approaches, in the hands of a practitioner, for detecting a model's reliance on spurious signals. 1 It is hard to find a needle in a haystack, it is much harder if you haven't seen a needle before (Pearl). -Judea Pearl
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Spurious Features Everywhere - Large-Scale Detection of Harmful Spurious Features in ImageNetYannic Neuhaus, Maximilian Augustin, Valentyn Boreiko, Matthias HeinICCV 2023 · 被引用 42 次
- "Will You Find These Shortcuts?" A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text ClassificationJasmijn Bastings, Sebastian Ebert, Polina Zablotskaia, Anders Sandholm 等EMNLP 2022 · 被引用 29 次
- Studying How to Efficiently and Effectively Guide Models with ExplanationsSukrut Rao, Moritz Böhle, Amin Parchami-Araghi, Bernt SchieleICCV 2023 · 被引用 22 次
- On the Relationship Between Explanation and Prediction: A Causal ViewAmir-Hossein Karimi, Krikamol Muandet, Simon Kornblith, Bernhard Schölkopf 等ICML 2023 · 被引用 20 次
- Stability Guarantees for Feature Attributions with Multiplicative SmoothingAnton Xue, Rajeev Alur, Eric WongNeurIPS 2023 · 被引用 18 次
它引用的顶会 Paper18
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann 等ICML 2020 · 被引用 1,233 次
- Manipulating and Measuring Model InterpretabilityForough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan 等CHI 2021 · 被引用 663 次
- Noise or Signal: The Role of Image Backgrounds in Object RecognitionKai Yuanqing Xiao, Logan Engstrom, Andrew Ilyas, Aleksander MadryICLR 2021 · 被引用 451 次
- An Investigation of Why Overparameterization Exacerbates Spurious CorrelationsShiori Sagawa, Aditi Raghunathan, Pang Wei Koh, Percy LiangICML 2020 · 被引用 436 次
- On Completeness-aware Concept-Based Explanations in Deep Neural NetworksChih-Kuan Yeh, Been Kim, Sercan Ömer Arik, Chun-Liang Li 等NeurIPS 2020 · 被引用 390 次
相关 Paper
- Debugging Tests for Model ExplanationsJulius Adebayo, Michael Muelly, Ilaria Liccardi, Been KimNeurIPS 2020 · 被引用 209 次
- An Empirical Study on Explanations in Out-of-Domain SettingsGeorge Chrysostomou, Nikolaos AletrasACL 2022
- Rethinking Explanation Evaluation Under the Retraining SchemeYi Cai, Thibaud Ardoin, Mayank Gulati, Gerhard WunderAAAI 2026
- Are All Spurious Features in Natural Language Alike? An Analysis through a Causal LensNitish Joshi, Xiang Pan, He HeEMNLP 2022 · 被引用 19 次
- Spuriosity Rankings: Sorting Data to Measure and Mitigate BiasesMazda Moayeri, Wenxiao Wang, Sahil Singla, Soheil FeiziNeurIPS 2023 · 被引用 19 次
