Understanding the role of importance weighting for deep learning
Da Xu, Yuting Ye, Chuanwei Ruan
Abstract
The recent paper by Byrd & Lipton (2019) , based on empirical observations, raises a major concern on the impact of importance weighting for the over-parameterized deep learning models. They observe that as long as the model can separate the training data, the impact of importance weighting diminishes as the training proceeds. Nevertheless, there lacks a rigorous characterization of this phenomenon. In this paper, we provide formal characterizations and theoretical justifications on the role of importance weighting with respect to the implicit bias of gradient descent and margin-based learning theory. We reveal both the optimization dynamics and generalization performance under deep learning models. Our work not only explains the various novel phenomenons observed for importance weighting in deep learning, but also extends to the studies where the weights are being optimized as part of the model, which applies to a number of topics under active research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 88a2016f-cc4b-461c-9890-4af5c53bbfa0Cited by top-tier papers14
- Self-supervised Learning is More Robust to Dataset ImbalanceHong Liu, Jeff Z. HaoChen, Adrien Gaidon, Tengyu MaICLR 2022 · 190 citations
- Balanced MSE for Imbalanced Visual RegressionJiawei Ren, Mingyuan Zhang, Cunjun Yu, Ziwei LiuCVPR 2022 · 163 citations
- Robust Learning with Progressive Data Expansion Against Spurious CorrelationYihe Deng, Yu Yang, Baharan Mirzasoleiman, Quanquan GuNeurIPS 2023 · 53 citations
- UMIX: Improving Importance Weighting for Subpopulation Shift via Uncertainty-Aware MixupZongbo Han, Zhipeng Liang, Fan Yang, Liu Liu et al.NeurIPS 2022 · 53 citations
- RLSbench: Domain Adaptation Under Relaxed Label ShiftSaurabh Garg, Nick Erickson, James Sharpnack, Alex Smola et al.ICML 2023 · 44 citations
Builds on3
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
- Rethinking Importance Weighting for Deep Learning under Distribution ShiftTongtong Fang, Nan Lu, Gang Niu, Masashi SugiyamaNeurIPS 2020 · 179 citations
- Adversarial Counterfactual Learning and Evaluation for Recommender SystemDa Xu, Chuanwei Ruan, Evren Körpeoglu, Sushant Kumar et al.NeurIPS 2020 · 36 citations
Related papers
- Mirror Descent Maximizes Generalized Margin and Can Be Implemented EfficientlyHaoyuan Sun, Kwangjun Ahn, Christos Thrampoulidis, Navid AzizanNeurIPS 2022 · 33 citations
- Conflicting Biases at the Edge of Stability: Norm versus Sharpness RegularizationMaria Matveev, Vit Fojtik, Hung-Hsu Chou, Gitta Kutyniok et al.ICML 2026
- Bias of Stochastic Gradient Descent or the Architecture: Disentangling the Effects of Overparameterization of Neural NetworksAmit Peleg, Matthias HeinICML 2024
- Towards understanding how momentum improves generalization in deep learningSamy Jelassi, Yuanzhi LiICML 2022 · 53 citations
- Why Do We Need Weight Decay in Modern Deep Learning?Francesco D'Angelo, Maksym Andriushchenko, Aditya Vardhan Varre, Nicolas FlammarionNeurIPS 2024 · 101 citations
