Is Importance Weighting Incompatible with Interpolating Classifiers?
Ke Alexander Wang, Niladri Shekhar Chatterji, Saminul Haque, Tatsunori Hashimoto
Abstract
Importance weighting is a classic technique to handle distribution shifts. However, prior work has presented strong empirical and theoretical evidence demonstrating that importance weights can have little to no effect on overparameterized neural networks. Is importance weighting truly incompatible with the training of overparameterized neural networks? Our paper answers this in the negative. We show that importance weighting fails not because of the overparameterization, but instead, as a result of using exponentially-tailed losses like the logistic or cross-entropy loss. As a remedy, we show that polynomially-tailed losses restore the effects of importance reweighting in correcting distribution shift in overparameterized models. We characterize the behavior of gradient descent on importance weighted polynomially-tailed losses with overparameterized linear models, and theoretically demonstrate the advantage of using polynomially-tailed losses in a label shift setting. Surprisingly, our theory shows that using weights that are obtained by exponentiating the classical unbiased importance weights can improve performance. Finally, we demonstrate the practical value of our analysis with neural network experiments on a subpopulation shift and a label shift dataset. When reweighted, our loss function can outperform reweighted cross-entropy by as much as 9% in test accuracy. Our loss function also gives test accuracies comparable to, or even exceeding, well-tuned state-of-the-art methods for correcting distribution shifts. * Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dffda5f9-6378-44cf-92fc-e75791280390Cited by top-tier papers3
- Distributionally Robust Optimization via Ball Oracle AccelerationYair Carmon, Danielle HauslerNeurIPS 2022 · 23 citations
- Minimum-Norm Interpolation Under Covariate ShiftNeil Mallinar, Austin Zane, Spencer Frei, Bin YuICML 2024 · 13 citations
- Understanding Why Generalized Reweighting Does Not Improve Over ERMRuntian Zhai, Chen Dan, J. Zico Kolter, Pradeep Kumar RavikumarICLR 2023 · 6 citations
Builds on6
- Long-tail learning via logit adjustmentAditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu Jain et al.ICLR 2021 · 937 citations
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 226 citations
- Label-Imbalanced and Group-Sensitive Classification under OverparameterizationGanesh Ramachandra Kini, Orestis Paraskevas, Samet Oymak, Christos ThrampoulidisNeurIPS 2021 · 122 citations
- Risk Bounds for Over-parameterized Maximum Margin Classification on Sub-Gaussian MixturesYuan Cao, Quanquan Gu, Mikhail BelkinNeurIPS 2021 · 57 citations
Related papers
- Optimizing importance weighting in the presence of sub-population shiftsFloris Holstege, Bram Wouters, Noud P. A. van Giersbergen, Cees G. H. DiksICLR 2025
- On Statistical Bias In Active Learning: How and When to Fix ItSebastian Farquhar, Yarin Gal, Tom RainforthICLR 2021 · 96 citations
- UMIX: Improving Importance Weighting for Subpopulation Shift via Uncertainty-Aware MixupZongbo Han, Zhipeng Liang, Fan Yang, Liu Liu et al.NeurIPS 2022 · 53 citations
- Rethinking Importance Weighting for Deep Learning under Distribution ShiftTongtong Fang, Nan Lu, Gang Niu, Masashi SugiyamaNeurIPS 2020 · 179 citations
- Understanding new tasks through the lens of training data via exponential tiltingSubha Maity, Mikhail Yurochkin, Moulinath Banerjee, Yuekai SunICLR 2023 · 1 citation
