Posterior Re-calibration for Imbalanced Datasets
Junjiao Tian, Yen-Cheng Liu, Nathaniel Glaser, Yen-Chang Hsu, Zsolt Kira
摘要
Neural Networks can perform poorly when the training label distribution is heavily imbalanced, as well as when the testing data differs from the training distribution. In order to deal with shift in the testing label distribution, which imbalance causes, we motivate the problem from the perspective of an optimal Bayes classifier and derive a post-training prior rebalancing technique that can be solved through a KL-divergence based optimization. This method allows a flexible post-training hyper-parameter to be efficiently tuned on a validation set and effectively modify the classifier margin to deal with this imbalance. We further combine this method with existing likelihood shift methods, re-interpreting them from the same Bayesian perspective, and demonstrating that our method can deal with both problems in a unified way. The resulting algorithm can be conveniently used on probabilistic classification problems agnostic to underlying architectures. Our results on six different datasets and five different architectures show state of art accuracy, including on large-scale imbalanced datasets such as iNaturalist for classification and Synthia for semantic segmentation. Please see https://github.com/GT-RIPL/UNO-IC.git for implementation. Related Work In this paper, we focus on prior (label) distribution shift resulting from various degrees of imbalance in the training label distributions, and testing on a different distribution or emphasizing a different distribution in the testing metrics (e.g. class-averaged accuracy). We therefore summarize methods for dealing with such imbalance. Imbalance -Data Level Methods The simplest methods for dealing with label imbalance during training are random under-sampling (RUS) which discards samples from the majority classes and random over-sampling (ROS) which re-samples from the minority classes [3] . While ROS is infeasible when the data imbalance is extreme, RUS tends to overfit the minority classes [4] . Synthetic
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper32
- Self-Supervised Aggregation of Diverse Experts for Test-Agnostic Long-Tailed RecognitionYifan Zhang, Bryan Hooi, Lanqing Hong, Jiashi FengNeurIPS 2022 · 被引用 214 次
- Self-supervised Learning is More Robust to Dataset ImbalanceHong Liu, Jeff Z. HaoChen, Adrien Gaidon, Tengyu MaICLR 2022 · 被引用 190 次
- Balanced MSE for Imbalanced Visual RegressionJiawei Ren, Mingyuan Zhang, Cunjun Yu, Ziwei LiuCVPR 2022 · 被引用 163 次
- TAM: Topology-Aware Margin Loss for Class-Imbalanced Node ClassificationJaeyun Song, Joonhyung Park, Eunho YangICML 2022 · 被引用 90 次
- RankSim: Ranking Similarity Regularization for Deep Imbalanced RegressionYu Gong, Greg Mori, Frederick TungICML 2022 · 被引用 68 次
它引用的顶会 Paper1
相关 Paper
- CDMAD: Class-Distribution-Mismatch-Aware Debiasing for Class-Imbalanced Semi-Supervised LearningHyuck Lee, Heeyoung KimCVPR 2024
- AUC Maximization under Positive Distribution ShiftAtsutoshi Kumagai, Tomoharu Iwata, Hiroshi Takahashi, Taishi Nishiyama 等NeurIPS 2024 · 被引用 7 次
- Open-Sampling: Exploring Out-of-Distribution data for Re-balancing Long-tailed datasetsHongxin Wei, Lue Tao, Renchunzi Xie, Lei Feng 等ICML 2022 · 被引用 46 次
- Towards Calibrated Model for Long-Tailed Visual Recognition from Prior PerspectiveZhengzhuo Xu, Zenghao Chai, Chun YuanNeurIPS 2021 · 被引用 77 次
- Maximum Likelihood with Bias-Corrected Calibration is Hard-To-Beat at Label Shift AdaptationAmr Alexandari, Anshul Kundaje, Avanti ShrikumarICML 2020 · 被引用 123 次
