Chroma-VAE: Mitigating Shortcut Learning with Generative Classifiers
Wanqian Yang, Polina Kirichenko, Micah Goldblum, Andrew Gordon Wilson
Abstract
Deep neural networks are susceptible to shortcut learning, using simple features to achieve low training loss without discovering essential semantic structure. Contrary to prior belief, we show that generative models alone are not sufficient to prevent shortcut learning, despite an incentive to recover a more comprehensive representation of the data than discriminative approaches. However, we observe that shortcuts are preferentially encoded with minimal information, a fact that generative models can exploit to mitigate shortcut learning. In particular, we propose Chroma-VAE, a two-pronged approach where a VAE classifier is initially trained to isolate the shortcut in a small latent subspace, allowing a secondary classifier to be trained on the complementary, shortcut-free latent subspace. In addition to demonstrating the efficacy of Chroma-VAE on benchmark and real-world shortcut learning tasks, our work highlights the potential for manipulating the latent space of generative classifiers to isolate or interpret specific correlations. 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Simple and Fast Group Robustness by Automatic Feature ReweightingShikai Qiu, Andres Potapczynski, Pavel Izmailov, Andrew Gordon WilsonICML 2023 · 78 citations
- Task Confusion and Catastrophic Forgetting in Class-Incremental Learning: A Mathematical Framework for Discriminative and Generative ModelingsMilad Khademi Nori, Il-Min KimNeurIPS 2024 · 6 citations
- Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward ModelingZhibin Duan, Guowei Rong, Zhuo Li, Bo Chen et al.ICML 2026 · 4 citations
- Decompose-and-Compose: A Compositional Approach to Mitigating Spurious CorrelationFahimeh Hosseini Noohdani, Parsa Hosseini, Aryan Yazdan Parast, Hamidreza Yaghoubi Araghi et al.CVPR 2024 · 3 citations
- The Devil Is in the Details: Tackling Unimodal Spurious Correlations for Generalizable Multimodal Reward ModelsZichao Li, Xueru Wen, Jie Lou, Yuqiu Ji et al.ICML 2025
Builds on10
- Just Train Twice: Improving Group Robustness without Training Group InformationEvan Zheran Liu, Behzad Haghgoo, Annie S. Chen, Aditi Raghunathan et al.ICML 2021 · 683 citations
- Environment Inference for Invariant LearningElliot Creager, Jörn-Henrik Jacobsen, Richard S. ZemelICML 2021 · 454 citations
- Noise or Signal: The Role of Image Backgrounds in Object RecognitionKai Yuanqing Xiao, Logan Engstrom, Andrew Ilyas, Aleksander MadryICLR 2021 · 451 citations
- An Investigation of Why Overparameterization Exacerbates Spurious CorrelationsShiori Sagawa, Aditi Raghunathan, Pang Wei Koh, Percy LiangICML 2020 · 436 citations
- No Subclass Left Behind: Fine-Grained Robustness in Coarse-Grained Classification ProblemsNimit Sharad Sohoni, Jared Dunnmon, Geoffrey Angus, Albert Gu et al.NeurIPS 2020 · 316 citations
Related papers
- On the Foundations of Shortcut LearningKatherine L. Hermann, Hossein Mobahi, Thomas Fel, Michael Curtis MozerICLR 2024 · 72 citations
- COMI: COrrect and MItigate Shortcut Learning Behavior in Deep Neural NetworksLili Zhao, Qi Liu, Linan Yue, Wei Chen et al.SIGIR 2024 · 9 citations
- Learning Concept Credible Models for Mitigating ShortcutsJiaxuan Wang, Sarah Jabbour, Maggie Makar, Michael W. Sjoding et al.NeurIPS 2022 · 8 citations
- Navigate Beyond Shortcuts: Debiased Learning through the Lens of Neural CollapseYining Wang, Junjie Sun, Chenyue Wang, Mi Zhang et al.CVPR 2024 · 6 citations
- Efficient Unsupervised Shortcut Learning Detection and Mitigation in TransformersLukas Kuhn, Sari Sadiya, Jörg Schlötterer, Florian Buettner et al.ICCV 2025 · 1 citation
