Saliency is a Possible Red Herring When Diagnosing Poor Generalization
Joseph D. Viviano, Becks Simpson, Francis Dutil, Yoshua Bengio, Joseph Paul Cohen
Abstract
Poor generalization is one symptom of models that learn to predict target variables using spuriously-correlated image features present only in the training distribution instead of the true image features that denote a class. It is often thought that this can be diagnosed visually using attribution (aka saliency) maps. We study if this assumption is correct. In some prediction tasks, such as for medical images, one may have some images with masks drawn by a human expert, indicating a region of the image containing relevant information to make the prediction. We study multiple methods that take advantage of such auxiliary labels, by training networks to ignore distracting features which may be found outside of the region of interest. This mask information is only used during training and has an impact on generalization accuracy depending on the severity of the shift between the training and test distributions. Surprisingly, while these methods improve generalization performance in the presence of a covariate shift, there is no strong correspondence between the correction of attribution towards the features a human expert has labelled as important and generalization performance. These results suggest that the root cause of poor generalization may not always be spatially defined, and raise questions about the utility of masks as "attribution priors" as well as saliency maps for explainable predictions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Diagnosing failures of fairness transfer across distribution shift in real-world medical settingsJessica Schrouff, Natalie Harris, Sanmi Koyejo, Ibrahim M. Alabdulmohsin et al.NeurIPS 2022 · 84 citations
- Overinterpretation reveals image classification model pathologiesBrandon Carter, Siddhartha Jain, Jonas Mueller, David GiffordNeurIPS 2021 · 59 citations
- Optimizing Relevance Maps of Vision Transformers Improves RobustnessHila Chefer, Idan Schwartz, Lior WolfNeurIPS 2022 · 55 citations
- Use-Case-Grounded Simulations for Explanation EvaluationValerie Chen, Nari Johnson, Nicholay Topin, Gregory Plumb et al.NeurIPS 2022 · 26 citations
- ACAT: Adversarial Counterfactual Attention for Classification and Detection in Medical ImagingAlessandro Fontanella, Antreas Antoniou, Wenwen Li, Joanna M. Wardlaw et al.ICML 2023 · 15 citations
Builds on3
- Interpretations are Useful: Penalizing Explanations to Align Neural Networks with Prior KnowledgeLaura Rieger, Chandan Singh, W. James Murdoch, Bin YuICML 2020 · 249 citations
- Learning explanations that are hard to varyGiambattista Parascandolo, Alexander Neitz, Antonio Orvieto, Luigi Gresele et al.ICLR 2021 · 221 citations
- What shapes feature representations? Exploring datasets, architectures, and trainingKatherine L. Hermann, Andrew K. LampinenNeurIPS 2020 · 186 citations
Related papers
- What You See is What You Classify: Black Box AttributionsSteven Stalder, Nathanaël Perraudin, Radhakrishna Achanta, Fernando Pérez-Cruz et al.NeurIPS 2022 · 15 citations
- Towards Robust Classification Model by Counterfactual and Invariant Data GenerationChun-Hao Chang, George-Alexandru Adam, Anna GoldenbergCVPR 2021
- Salient ImageNet: How to discover spurious features in Deep Learning?Sahil Singla, Soheil FeiziICLR 2022 · 144 citations
- Causally motivated multi-shortcut identification and removalJiayun Zheng, Maggie MakarNeurIPS 2022 · 27 citations
- From Attribution to Action: Jointly ALIGNing Predictions and ExplanationsDongsheng Hong, Chao Chen, Yanhui Chen, Shanshan Lin et al.AAAI 2026
