Balancing the Scales: A Theoretical and Algorithmic Framework for Learning from Imbalanced Data
Corinna Cortes, Anqi Mao, Mehryar Mohri, Yutao Zhong
摘要
Class imbalance remains a major challenge in machine learning, especially in multi-class problems with long-tailed distributions. Existing methods, such as data resampling, cost-sensitive techniques, and logistic loss modifications, though popular and often effective, lack solid theoretical foundations. As an example, we demonstrate that cost-sensitive methods are not Bayes consistent. This paper introduces a novel theoretical framework for analyzing generalization in imbalanced classification. We propose a new class-imbalanced margin loss function for both binary and multiclass settings, prove its strong H-consistency, and derive corresponding learning guarantees based on empirical loss and a new notion of class-sensitive Rademacher complexity. Leveraging these theoretical results, we devise novel and general learning algorithms, IMMAX (Imbalanced Margin Maximization), which incorporate confidence margins and are applicable to various hypothesis sets. While our focus is theoretical, we also present extensive empirical results demonstrating the effectiveness of our algorithms compared to existing baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Improved Balanced Classification with Theoretically Grounded Loss FunctionsCorinna Cortes, Mehryar Mohri, Yutao ZhongNeurIPS 2025 · 被引用 19 次
- Why Ask One When You Can Ask k? Learning-to-Defer to the Top-k ExpertsYannis Montreuil, Axel Carlier, Lai Xing Ng, Wei Tsang OoiICLR 2026 · 被引用 7 次
- Linear-Core Surrogates: Smooth Loss Functions with Linear Rates for Classification and Structured PredictionMehryar Mohri, Yutao ZhongICML 2026 · 被引用 7 次
- Optimized Deferral for Imbalanced SettingsCorinna Cortes, Anqi Mao, Mehryar Mohri, Yutao ZhongICML 2026 · 被引用 7 次
- Reducing Class-Wise Performance Disparity via Margin RegularizationBeier Zhu, Kesen Zhao, Jiequan Cui, Qianru Sun 等ICLR 2026 · 被引用 1 次
它引用的顶会 Paper52
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan 等ICLR 2020 · 被引用 1,496 次
- Long-tail learning via logit adjustmentAditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu Jain 等ICLR 2021 · 被引用 937 次
- Balanced Meta-Softmax for Long-Tailed Visual RecognitionJiawei Ren, Cunjun Yu, Shunan Sheng, Xiao Ma 等NeurIPS 2020 · 被引用 861 次
- Cross-Entropy Loss Functions: Theoretical Analysis and ApplicationsAnqi Mao, Mehryar Mohri, Yutao ZhongICML 2023 · 被引用 790 次
- Long-Tailed Classification by Keeping the Good and Removing the Bad Momentum Causal EffectKaihua Tang, Jianqiang Huang, Hanwang ZhangNeurIPS 2020 · 被引用 533 次
相关 Paper
- Principled Algorithms for Optimizing Generalized Metrics in Binary ClassificationAnqi Mao, Mehryar Mohri, Yutao ZhongICML 2025
- Multi-Class -Consistency BoundsPranjal Awasthi, Anqi Mao, Mehryar Mohri, Yutao ZhongNeurIPS 2022 · 被引用 48 次
- Escaping Saddle Points for Effective Generalization on Class-Imbalanced DataHarsh Rangwani, Sumukh K. Aithal, Mayank Mishra, Venkatesh Babu R.NeurIPS 2022 · 被引用 50 次
- Difficulty-aware Balancing Margin Loss for Long-tailed RecognitionMinseok Son, Inyong Koo, Jinyoung Park, Changick KimAAAI 2025 · 被引用 9 次
- Relative Deviation Margin BoundsCorinna Cortes, Mehryar Mohri, Ananda Theertha SureshICML 2021 · 被引用 16 次
