DiagViB-6: A Diagnostic Benchmark Suite for Vision Models in the Presence of Shortcut and Generalization Opportunities
Elias Eulig, Piyapat Saranrittichai, Chaithanya Kumar Mummadi, Kilian Rambach, William Beluch, Xiahan Shi, Volker Fischer
摘要
Common deep neural networks (DNNs) for image classification have been shown to rely on shortcut opportunities (SO) in the form of predictive and easy-to-represent visual factors. This is known as shortcut learning and leads to impaired generalization. In this work, we show that common DNNs also suffer from shortcut learning when predicting only basic visual object factors of variation (FoV) such as shape, color, or texture. We argue that besides shortcut opportunities, generalization opportunities (GO) are also an inherent part of real-world vision data and arise from partial independence between predicted classes and FoVs. We also argue that it is necessary for DNNs to exploit GO to overcome shortcut learning. Our core contribution is to introduce the Diagnostic Vision Benchmark suite DiagViB-6, which includes datasets and metrics to study a network’s shortcut vulnerability and generalization capability for six independent FoV. In particular, DiagViB-6 allows controlling the type and degree of SO and GO in a dataset. We benchmark a wide range of popular vision architectures and show that they can exploit GO only to a limited extent.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Visual Representation Learning Does Not Generalize Strongly Within the Same DomainLukas Schott, Julius von Kügelgen, Frederik Träuble, Peter Vincent Gehler 等ICLR 2022 · 被引用 79 次
- Which Shortcut Cues Will DNNs Choose? A Study from the Parameter-Space PerspectiveLuca Scimeca, Seong Joon Oh, Sanghyuk Chun, Michael Poli 等ICLR 2022 · 被引用 67 次
- Invariant Anomaly Detection under Distribution Shifts: A Causal PerspectiveJoão B. S. Carvalho, Mengtao Zhang, Robin Geyer, Carlos Cotrini 等NeurIPS 2023 · 被引用 18 次
- A Whac-A-Mole Dilemma: Shortcuts Come in Multiples Where Mitigating One Amplifies OthersZhiheng Li, Ivan Evtimov, Albert Gordo, Caner Hazirbas 等CVPR 2023
它引用的顶会 Paper8
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 被引用 1,188 次
- Unsupervised Discovery of Interpretable Directions in the GAN Latent SpaceAndrey Voynov, Artem BabenkoICML 2020 · 被引用 459 次
- The Origins and Prevalence of Texture Bias in Convolutional Neural NetworksKatherine L. Hermann, Ting Chen, Simon KornblithNeurIPS 2020 · 被引用 369 次
- What shapes feature representations? Exploring datasets, architectures, and trainingKatherine L. Hermann, Andrew K. LampinenNeurIPS 2020 · 被引用 186 次
- Counterfactual Generative NetworksAxel Sauer, Andreas GeigerICLR 2021 · 被引用 145 次
相关 Paper
- A New Benchmark: On the Utility of Synthetic Data with Blender for Bare Supervised Learning and Downstream Domain AdaptationHui Tang, Kui JiaCVPR 2023
- On the Foundations of Shortcut LearningKatherine L. Hermann, Hossein Mobahi, Thomas Fel, Michael Curtis MozerICLR 2024 · 被引用 72 次
- Do ImageNet-trained Models Learn Shortcuts? The Impact of Frequency Shortcuts on GeneralizationShunxin Wang, Raymond N. J. Veldhuis, Nicola StrisciuglioCVPR 2025
- What do neural networks learn in image classification? A frequency shortcut perspectiveShunxin Wang, Raymond N. J. Veldhuis, Christoph Brune, Nicola StrisciuglioICCV 2023 · 被引用 51 次
- Efficient Unsupervised Shortcut Learning Detection and Mitigation in TransformersLukas Kuhn, Sari Sadiya, Jörg Schlötterer, Florian Buettner 等ICCV 2025 · 被引用 1 次
