Salient ImageNet: How to discover spurious features in Deep Learning?
Sahil Singla, Soheil Feizi
Abstract
we identify spurious or core neural features (penultimate layer neurons of a robust model) via limited human supervision (e.g., using top 5 activating images per feature). We then show that these neural feature annotations generalize extremely well to many more images without any human supervision. We use the activation maps for these neural features as the soft masks to highlight spurious or core visual features. Using this methodology, we introduce the Salient Imagenet dataset containing core and spurious masks for a large set of samples from Imagenet. Using this dataset, we show that several popular Imagenet models rely heavily on various spurious features in their predictions, in-dicating the standard accuracy alone is not sufficient to fully assess model perfor-1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ce061ea-6450-442b-a29b-216138029072Cited by top-tier papers47
- Entropy is not Enough for Test-Time Adaptation: From the Perspective of Disentangled FactorsJonghyun Lee, Dahuin Jung, Saehyung Lee, Junsung Park et al.ICLR 2024 · 106 citations
- Discover and Cure: Concept-aware Mitigation of Spurious CorrelationShirley Wu, Mert Yüksekgönül, Linjun Zhang, James ZouICML 2023 · 97 citations
- Towards Last-layer Retraining for Group Robustness with Fewer AnnotationsTyler LaBonte, Vidya Muthukumar, Abhishek KumarNeurIPS 2023 · 73 citations
- On the Foundations of Shortcut LearningKatherine L. Hermann, Hossein Mobahi, Thomas Fel, Michael Curtis MozerICLR 2024 · 72 citations
- LANCE: Stress-testing Visual Models by Generating Language-guided Counterfactual ImagesViraj Prabhu, Sriram Yenamandra, Prithvijit Chattopadhyay, Judy HoffmanNeurIPS 2023 · 59 citations
Builds on8
- Noise or Signal: The Role of Image Backgrounds in Object RecognitionKai Yuanqing Xiao, Logan Engstrom, Andrew Ilyas, Aleksander MadryICLR 2021 · 451 citations
- Benchmarking Deep Learning Interpretability in Time Series PredictionsAya Abdelsalam Ismail, Mohamed K. Gunady, Héctor Corrada Bravo, Soheil FeiziNeurIPS 2020 · 249 citations
- From ImageNet to Image Classification: Contextualizing Progress on BenchmarksDimitris Tsipras, Shibani Santurkar, Logan Engstrom, Andrew Ilyas et al.ICML 2020 · 146 citations
- Leveraging Sparse Linear Layers for Debuggable Deep NetworksEric Wong, Shibani Santurkar, Aleksander MadryICML 2021 · 101 citations
- Generative causal explanations of black-box classifiersMatthew R. O'Shaughnessy, Gregory Canal, Marissa Connor, Christopher Rozell et al.NeurIPS 2020 · 83 citations
Related papers
- Spurious Features Everywhere - Large-Scale Detection of Harmful Spurious Features in ImageNetYannic Neuhaus, Maximilian Augustin, Valentyn Boreiko, Matthias HeinICCV 2023 · 42 citations
- Spuriosity Rankings: Sorting Data to Measure and Mitigate BiasesMazda Moayeri, Wenxiao Wang, Sahil Singla, Soheil FeiziNeurIPS 2023 · 19 citations
- Towards Robust Classification Model by Counterfactual and Invariant Data GenerationChun-Hao Chang, George-Alexandru Adam, Anna GoldenbergCVPR 2021
- Last Layer Re-Training is Sufficient for Robustness to Spurious CorrelationsPolina Kirichenko, Pavel Izmailov, Andrew Gordon WilsonICLR 2023 · 31 citations
- Saliency is a Possible Red Herring When Diagnosing Poor GeneralizationJoseph D. Viviano, Becks Simpson, Francis Dutil, Yoshua Bengio et al.ICLR 2021 · 46 citations
