Don't blame Dataset Shift! Shortcut Learning due to Gradients and Cross Entropy
Aahlad Manas Puli, Lily H. Zhang, Yoav Wald, Rajesh Ranganath
摘要
Common explanations for shortcut learning assume that the shortcut improves prediction under the training distribution but not in the test distribution. Thus, models trained via the typical gradient-based optimization of cross-entropy, which we call default-ERM, utilize the shortcut. However, even when the stable feature determines the label in the training distribution and the shortcut does not provide any additional information, like in perception tasks, default-ERM still exhibits shortcut learning. Why are such solutions preferred when the loss for default-ERM can be driven to zero using the stable feature alone? By studying a linear perception task, we show that default-ERM's preference for maximizing the margin leads to models that depend more on the shortcut than the stable feature, even without overparameterization. This insight suggests that default-ERM's implicit inductive bias towards max-margin is unsuitable for perception tasks. Instead, we develop an inductive bias toward uniform margins and show that this bias guarantees dependence only on the perfect stable feature in the linear perception task. We develop loss functions that encourage uniform-margin solutions, called margin control (MARG-CTRL). MARG-CTRL mitigates shortcut learning on a variety of vision and language tasks, showing that better inductive biases can remove the need for expensive two-stage shortcut-mitigating methods in perception tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Enhancing Domain Adaptation through Prompt Gradient AlignmentViet Hoang Phan, Tung Lam Tran, Quyen Tran, Trung LeNeurIPS 2024 · 被引用 18 次
- Improving Subgroup Robustness via Data SelectionSaachi Jain, Kimia Hamidieh, Kristian Georgiev, Andrew Ilyas 等NeurIPS 2024 · 被引用 17 次
- Causal-structure Driven Augmentations for Text OOD GeneralizationAmir Feder, Yoav Wald, Claudia Shi, Suchi Saria 等NeurIPS 2023 · 被引用 10 次
- Changing the Training Data Distribution to Reduce Simplicity Bias Improves In-distribution GeneralizationDang Nguyen, Paymon Haddad, Eric Gan, Baharan MirzasoleimanNeurIPS 2024 · 被引用 4 次
- Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across TasksDamien Teney, Liangze Jiang, Hemanth Saratchandran, Simon LuceyICLR 2026 · 被引用 3 次
它引用的顶会 Paper23
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 被引用 1,578 次
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 被引用 1,416 次
- Out-of-Distribution Generalization via Risk Extrapolation (REx)David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang 等ICML 2021 · 被引用 1,163 次
- Just Train Twice: Improving Group Robustness without Training Group InformationEvan Zheran Liu, Behzad Haghgoo, Annie S. Chen, Aditi Raghunathan 等ICML 2021 · 被引用 683 次
相关 Paper
- COMI: COrrect and MItigate Shortcut Learning Behavior in Deep Neural NetworksLili Zhao, Qi Liu, Linan Yue, Wei Chen 等SIGIR 2024 · 被引用 9 次
- Mitigating Shortcut Learning with InterpoLated LearningMichalis Korakakis, Andreas Vlachos, Adrian WellerACL 2025
- On the Foundations of Shortcut LearningKatherine L. Hermann, Hossein Mobahi, Thomas Fel, Michael Curtis MozerICLR 2024 · 被引用 72 次
- Shortcut Features as Top Eigenfunctions of NTK: A Linear Neural Network Case and MoreJinwoo Lim, Suhyun Kim, Soo-Mook MoonNeurIPS 2025 · 被引用 1 次
- Which Shortcut Cues Will DNNs Choose? A Study from the Parameter-Space PerspectiveLuca Scimeca, Seong Joon Oh, Sanghyuk Chun, Michael Poli 等ICLR 2022 · 被引用 67 次
