On Harmonizing Implicit Subpopulations
Feng Hong, Jiangchao Yao, Yueming Lyu, Zhihan Zhou, Ivor W. Tsang, Ya Zhang, Yanfeng Wang
Abstract
Machine learning algorithms learned from data with skewed distributions usually suffer from poor generalization, especially when minority classes matter as much as, or even more than majority ones. This is more challenging on class-balanced data that has some hidden imbalanced subpopulations, since prevalent techniques mainly conduct class-level calibration and cannot perform subpopulation-level adjustments without subpopulation annotations. Regarding implicit subpopulation imbalance, we reveal that the key to alleviating the detrimental effect lies in effective subpopulation discovery with proper rebalancing. We then propose a novel subpopulation-imbalanced learning method called Scatter and HarmonizE (SHE). Our method is built upon the guiding principle of optimal data partition, which involves assigning data to subpopulations in a manner that maximizes the predictive information from inputs to labels. With theoretical guarantees and empirical evidences, SHE succeeds in identifying the hidden subpopulations and encourages subpopulation-balanced predictions. Extensive experiments on various benchmark datasets show the effectiveness of SHE. The code is available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Locally Estimated Global Perturbations are Better than Local Perturbations for Federated Sharpness-aware MinimizationZiqing Fan, Shengchao Hu, Jiangchao Yao, Gang Niu et al.ICML 2024 · 35 citations
- Revive Re-weighting in Imbalanced Learning by Density Ratio EstimationJiaan Luo, Feng Hong, Jiangchao Yao, Bo Han et al.NeurIPS 2024 · 16 citations
- Diversified Batch Selection for Training AccelerationFeng Hong, Yueming Lyu, Jiangchao Yao, Ya Zhang et al.ICML 2024 · 16 citations
- Long-tailed Recognition with Model RebalancingJiaan Luo, Feng Hong, Qiang Hu, Xiaofeng Cao et al.NeurIPS 2025 · 12 citations
- Mitigating Noisy Correspondence by Geometrical Structure Consistency LearningZihua Zhao, Mengxi Chen, Tianjie Dai, Jiangchao Yao et al.CVPR 2024 · 5 citations
Builds on49
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
Related papers
- Boosting Test Performance with Importance Sampling-a Subpopulation PerspectiveHongyu Shen, Zhizhen ZhaoAAAI 2025
- CReST: A Class-Rebalancing Self-Training Framework for Imbalanced Semi-Supervised LearningChen Wei, Kihyuk Sohn, Clayton Mellina, Alan L. Yuille et al.CVPR 2021
- Procrustean Training for Imbalanced Deep LearningHan-Jia Ye, De-Chuan Zhan, Wei-Lun ChaoICCV 2021 · 36 citations
- Self-paced Ensemble for Highly Imbalanced Massive Data ClassificationZhining Liu, Wei Cao, Zhifeng Gao, Jiang Bian et al.ICDE 2020 · 172 citations
- Representation Matters: Assessing the Importance of Subgroup Allocations in Training DataEsther Rolf, Theodora T. Worledge, Benjamin Recht, Michael I. JordanICML 2021 · 50 citations
