Boosting Test Performance with Importance Sampling-a Subpopulation Perspective
Hongyu Shen, Zhizhen Zhao
Abstract
Despite empirical risk minimization (ERM) is widely applied in the machine learning community, its performance is limited on data with spurious correlation or subpopulation that is introduced by hidden attributes. Existing literature proposed techniques to maximize group-balanced or worst-group accuracy when such correlation presents, yet, at the cost of lower average accuracy. In addition, many existing works conduct surveys on different subpopulation methods without revealing the inherent connection between these methods, which could hinder the technology advancement in this area. In this paper, we identify important sampling as a simple yet powerful tool for solving the subpopulation problem. On the theory side, we provide a new systematic formulation of the subpopulation problem, and explicitly identify the assumptions that are not clearly stated in the existing works. This helps to uncover the cause of the dropped average accuracy. We provide the first theoretical discussion on the connections of existing methods, revealing the core components that make them different. On the application side, we demonstrate a single estimator is enough to solve the subpopulation problem. In particular, we introduce the estimator in both attribute-known and -unknown scenarios in the subpopulation setup, offering flexibility in practical use cases. And empirically, we achieve state-of-the-art performance on commonly used benchmark datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 43674617-c8e8-4dec-8f69-2d293beac90aBuilds on15
- Long-tail learning via logit adjustmentAditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu Jain et al.ICLR 2021 · 937 citations
- Just Train Twice: Improving Group Robustness without Training Group InformationEvan Zheran Liu, Behzad Haghgoo, Annie S. Chen, Aditi Raghunathan et al.ICML 2021 · 683 citations
- Learning from Failure: De-biasing Classifier from Biased ClassifierJun Hyun Nam, Hyuntak Cha, Sungsoo Ahn, Jaeho Lee et al.NeurIPS 2020 · 428 citations
- Improving Out-of-Distribution Robustness via Selective AugmentationHuaxiu Yao, Yu Wang, Sai Li, Linjun Zhang et al.ICML 2022 · 275 citations
- On Feature Learning in the Presence of Spurious CorrelationsPavel Izmailov, Polina Kirichenko, Nate Gruver, Andrew Gordon WilsonNeurIPS 2022 · 208 citations
Related papers
- Gradient Extrapolation for Debiased Representation LearningIhab Asaad, Maha Shadaydeh, Joachim DenzlerICCV 2025 · 4 citations
- Correct-N-Contrast: a Contrastive Approach for Improving Robustness to Spurious CorrelationsMichael Zhang, Nimit Sharad Sohoni, Hongyang R. Zhang, Chelsea Finn et al.ICML 2022 · 230 citations
- Avoiding spurious correlations via logit correctionSheng Liu, Xu Zhang, Nitesh Sekhar, Yue Wu et al.ICLR 2023 · 3 citations
- Mitigating Spurious Correlations via Disagreement ProbabilityHyeonggeun Han, Sehwan Kim, Hyungjun Joo, Sangwoo Hong et al.NeurIPS 2024 · 8 citations
- When do Minimax-fair Learning and Empirical Risk Minimization Coincide?Harvineet Singh, Matthäus Kleindessner, Volkan Cevher, Rumi Chunara et al.ICML 2023 · 6 citations
