UMIX: Improving Importance Weighting for Subpopulation Shift via Uncertainty-Aware Mixup
Zongbo Han, Zhipeng Liang, Fan Yang, Liu Liu, Lanqing Li, Yatao Bian, Peilin Zhao, Bingzhe Wu, Changqing Zhang, Jianhua Yao
Abstract
Subpopulation shift widely exists in many real-world machine learning applications, referring to the training and test distributions containing the same subpopulation groups but varying in subpopulation frequencies. Importance reweighting is a normal way to handle the subpopulation shift issue by imposing constant or adaptive sampling weights on each sample in the training dataset. However, some recent studies have recognized that most of these approaches fail to improve the performance over empirical risk minimization especially when applied to over-parameterized neural networks. In this work, we propose a simple yet practical framework, called uncertainty-aware mixup (UMIX), to mitigate the overfitting issue in over-parameterized models by reweighting the ''mixed'' samples according to the sample uncertainty. The training-trajectories-based uncertainty estimation is equipped in the proposed UMIX for each sample to flexibly characterize the subpopulation distribution. We also provide insightful theoretical analysis to verify that UMIX achieves better generalization bounds over prior works. Further, we conduct extensive empirical studies across a wide range of tasks to validate the effectiveness of our method both qualitatively and quantitatively. Code is available at https://github.com/TencentAILabHealthcare/UMIX.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d992beb5-0021-4e9b-83e7-43a96c805a76Cited by top-tier papers22
- Change is Hard: A Closer Look at Subpopulation ShiftYuzhe Yang, Haoran Zhang, Dina Katabi, Marzyeh GhassemiICML 2023 · 149 citations
- Provable Dynamic Fusion for Low-Quality Multimodal DataQingyang Zhang, Haitao Wu, Changqing Zhang, Qinghua Hu et al.ICML 2023 · 143 citations
- Discover and Cure: Concept-aware Mitigation of Spurious CorrelationShirley Wu, Mert Yüksekgönül, Linjun Zhang, James ZouICML 2023 · 97 citations
- Enhancing Minority Classes by Mixing: An Adaptative Optimal Transport Approach for Long-tailed ClassificationJintong Gao, He Zhao, Zhuo Li, Dandan GuoNeurIPS 2023 · 64 citations
- Geodesic Multi-Modal Mixup for Robust Fine-TuningChangdae Oh, Junhyuk So, Hoyoon Byun, YongTaek Lim et al.NeurIPS 2023 · 49 citations
Builds on27
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
- Out-of-Distribution Generalization via Risk Extrapolation (REx)David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang et al.ICML 2021 · 1,163 citations
- Just Train Twice: Improving Group Robustness without Training Group InformationEvan Zheran Liu, Behzad Haghgoo, Annie S. Chen, Aditi Raghunathan et al.ICML 2021 · 683 citations
Related papers
- Optimizing importance weighting in the presence of sub-population shiftsFloris Holstege, Bram Wouters, Noud P. A. van Giersbergen, Cees G. H. DiksICLR 2025
- Diverse Prototypical Ensembles Improve Robustness to Subpopulation ShiftMinh Nguyen Nhat To, Paul F. R. Wilson, Viet Nguyen, Mohamed Harmanani et al.ICML 2025
- Understanding Why Generalized Reweighting Does Not Improve Over ERMRuntian Zhai, Chen Dan, J. Zico Kolter, Pradeep Kumar RavikumarICLR 2023 · 6 citations
- Is Importance Weighting Incompatible with Interpolating Classifiers?Ke Alexander Wang, Niladri Shekhar Chatterji, Saminul Haque, Tatsunori HashimotoICLR 2022 · 22 citations
- Overparameterisation and worst-case generalisation: friend or foe?Aditya Krishna Menon, Ankit Singh Rawat, Sanjiv KumarICLR 2021 · 38 citations
