Focus on the Common Good: Group Distributional Robustness Follows
Vihari Piratla, Praneeth Netrapalli, Sunita Sarawagi
Abstract
We consider the problem of training a classification model with group annotated training data. Recent work has established that, if there is distribution shift across different groups, models trained using the standard empirical risk minimization (ERM) objective suffer from poor performance on minority groups and that group distributionally robust optimization (Group-DRO) objective is a better alternative. The starting point of this paper is the observation that though Group-DRO performs better than ERM on minority groups for some benchmark datasets, there are several other datasets where it performs much worse than ERM. Inspired by ideas from the closely related problem of domain generalization, this paper proposes a new and simple algorithm that explicitly encourages learning of features that are shared across various groups. The key insight behind our proposed algorithm is that while Group-DRO focuses on groups with worst regularized loss, focusing instead, on groups that enable better performance even on other groups, could lead to learning of shared/common features, thereby enhancing minority performance beyond what is achieved by Group-DRO. Empirically, we show that our proposed algorithm matches or achieves better performance compared to strong contemporary baselines including ERM and Group-DRO on standard benchmarks on both minority groups and across all groups. Theoretically, we show that the proposed algorithm is a descent method and finds first order stationary points of smooth nonconvex functions. Our code and datasets can be found at this URL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b27d749b-2bb5-4c07-8e1c-4ab99d79494bCited by top-tier papers16
- UMIX: Improving Importance Weighting for Subpopulation Shift via Uncertainty-Aware MixupZongbo Han, Zhipeng Liang, Fan Yang, Liu Liu et al.NeurIPS 2022 · 53 citations
- Robust Multi-Task Learning with Excess RisksYifei He, Shiji Zhou, Guojun Zhang, Hyokun Yun et al.ICML 2024 · 29 citations
- Domain Generalization via Rationale InvarianceLiang Chen, Yong Zhang, Yibing Song, Anton van den Hengel et al.ICCV 2023 · 29 citations
- Intersectional Two-sided Fairness in RecommendationYifan Wang, Peijie Sun, Weizhi Ma, Min Zhang et al.WWW 2024 · 27 citations
- Temporally and Distributionally Robust Optimization for Cold-Start RecommendationXinyu Lin, Wenjie Wang, Jujia Zhao, Yongqi Li et al.AAAI 2024 · 23 citations
Builds on11
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
- Just Train Twice: Improving Group Robustness without Training Group InformationEvan Zheran Liu, Behzad Haghgoo, Annie S. Chen, Aditi Raghunathan et al.ICML 2021 · 683 citations
- Environment Inference for Invariant LearningElliot Creager, Jörn-Henrik Jacobsen, Richard S. ZemelICML 2021 · 454 citations
- An Investigation of Why Overparameterization Exacerbates Spurious CorrelationsShiori Sagawa, Aditi Raghunathan, Pang Wei Koh, Percy LiangICML 2020 · 436 citations
Related papers
- Examining and Combating Spurious Features under Distribution ShiftChunting Zhou, Xuezhe Ma, Paul Michel, Graham NeubigICML 2021 · 78 citations
- Re-weighting Based Group Fairness Regularization via Classwise Robust OptimizationSangwon Jung, Taeeon Park, Sanghyuk Chun, Taesup MoonICLR 2023 · 5 citations
- Understanding Why Generalized Reweighting Does Not Improve Over ERMRuntian Zhai, Chen Dan, J. Zico Kolter, Pradeep Kumar RavikumarICLR 2023 · 6 citations
- Mitigating Spurious Correlation via Distributionally Robust Learning with Hierarchical Ambiguity SetsSung Ho Jo, Seonghwi Kim, Minwoo ChaeICLR 2026 · 6 citations
- MixMax: Distributional Robustness in Function Space via Optimal Data MixturesAnvith Thudi, Chris J. MaddisonICLR 2025
