DiagViB-6: A Diagnostic Benchmark Suite for Vision Models in the Presence of Shortcut and Generalization Opportunities
Elias Eulig, Piyapat Saranrittichai, Chaithanya Kumar Mummadi, Kilian Rambach, William Beluch, Xiahan Shi, Volker Fischer
Abstract
Common deep neural networks (DNNs) for image classification have been shown to rely on shortcut opportunities (SO) in the form of predictive and easy-to-represent visual factors. This is known as shortcut learning and leads to impaired generalization. In this work, we show that common DNNs also suffer from shortcut learning when predicting only basic visual object factors of variation (FoV) such as shape, color, or texture. We argue that besides shortcut opportunities, generalization opportunities (GO) are also an inherent part of real-world vision data and arise from partial independence between predicted classes and FoVs. We also argue that it is necessary for DNNs to exploit GO to overcome shortcut learning. Our core contribution is to introduce the Diagnostic Vision Benchmark suite DiagViB-6, which includes datasets and metrics to study a network’s shortcut vulnerability and generalization capability for six independent FoV. In particular, DiagViB-6 allows controlling the type and degree of SO and GO in a dataset. We benchmark a wide range of popular vision architectures and show that they can exploit GO only to a limited extent.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ac9c1a68-8cfd-404b-b5e4-916ae79cbaceCited by top-tier papers4
- Visual Representation Learning Does Not Generalize Strongly Within the Same DomainLukas Schott, Julius von Kügelgen, Frederik Träuble, Peter Vincent Gehler et al.ICLR 2022 · 79 citations
- Which Shortcut Cues Will DNNs Choose? A Study from the Parameter-Space PerspectiveLuca Scimeca, Seong Joon Oh, Sanghyuk Chun, Michael Poli et al.ICLR 2022 · 67 citations
- Invariant Anomaly Detection under Distribution Shifts: A Causal PerspectiveJoão B. S. Carvalho, Mengtao Zhang, Robin Geyer, Carlos Cotrini et al.NeurIPS 2023 · 18 citations
- A Whac-A-Mole Dilemma: Shortcuts Come in Multiples Where Mitigating One Amplifies OthersZhiheng Li, Ivan Evtimov, Albert Gordo, Caner Hazirbas et al.CVPR 2023
Builds on8
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 1,188 citations
- Unsupervised Discovery of Interpretable Directions in the GAN Latent SpaceAndrey Voynov, Artem BabenkoICML 2020 · 459 citations
- The Origins and Prevalence of Texture Bias in Convolutional Neural NetworksKatherine L. Hermann, Ting Chen, Simon KornblithNeurIPS 2020 · 369 citations
- What shapes feature representations? Exploring datasets, architectures, and trainingKatherine L. Hermann, Andrew K. LampinenNeurIPS 2020 · 186 citations
- Counterfactual Generative NetworksAxel Sauer, Andreas GeigerICLR 2021 · 145 citations
Related papers
- A New Benchmark: On the Utility of Synthetic Data with Blender for Bare Supervised Learning and Downstream Domain AdaptationHui Tang, Kui JiaCVPR 2023
- On the Foundations of Shortcut LearningKatherine L. Hermann, Hossein Mobahi, Thomas Fel, Michael Curtis MozerICLR 2024 · 72 citations
- Do ImageNet-trained Models Learn Shortcuts? The Impact of Frequency Shortcuts on GeneralizationShunxin Wang, Raymond N. J. Veldhuis, Nicola StrisciuglioCVPR 2025
- What do neural networks learn in image classification? A frequency shortcut perspectiveShunxin Wang, Raymond N. J. Veldhuis, Christoph Brune, Nicola StrisciuglioICCV 2023 · 51 citations
- Efficient Unsupervised Shortcut Learning Detection and Mitigation in TransformersLukas Kuhn, Sari Sadiya, Jörg Schlötterer, Florian Buettner et al.ICCV 2025 · 1 citation
