Kill Two Birds with One Stone: Rethinking Data Augmentation for Deep Long-tailed Learning
Binwu Wang, Pengkun Wang, Wei Xu, Xu Wang, Yudong Zhang, Kun Wang, Yang Wang
Abstract
Real-world tasks are universally associated with training samples that exhibit a long-tailed class distribution, and traditional deep learning models are not suitable for fitting this distribution, thus resulting in a biased trained model. To surmount this dilemma, massive deep long-tailed learning studies have been proposed to achieve inter-class fairness models by designing sophisticated sampling strategies or improving existing model structures and loss functions. Habitually, these studies tend to apply data augmentation strategies to improve the generalization performance of their models. However, this augmentation strategy applied to balanced distributions may not be the best option for long-tailed distributions. For a profound understanding of data augmentation, we first theoretically analyze the gains of traditional augmentation strategies in long-tailed learning, and observe that augmentation methods cause the long-tailed distribution to be imbalanced again, resulting in an intertwined imbalance: inherent data-wise imbalance and extrinsic augmentation-wise imbalance, i.e., two 'birds' co-exist in long-tailed learning. Motivated by this observation, we propose an adaptive Dynamic Optional Data Augmentation (DODA) to address this intertwined imbalance, i.e., one 'stone' simultaneously 'kills' two 'birds', which allows each class to choose appropriate augmentation methods by maintaining a corresponding augmentation probability distribution for each class during training. Extensive experiments across mainstream long-tailed recognition benchmarks (e.g., CIFAR-100-LT, ImageNet-LT, and iNaturalist 2018) prove the effectiveness and flexibility of the DODA in overcoming the intertwined imbalance.
Recently, massive deep long-tailed learning studies have been proposed to surmount the class imbalance problem. The most intuitive and mainstream paradigm is class re-balancing, which balances the training sample numbers or weights of different classes during model training by resampling Kang et al. ( 2020); Ren et al. (2020); Wang et al. (2020); Jia et al. (2023) or cost-sensitive
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 99e5d5b5-5afc-48d0-962d-2488eb42053dCited by top-tier papers15
- LLM-AutoDA: Large Language Model-Driven Automatic Data Augmentation for Long-tailed ProblemsPengkun Wang, Zhe Zhao, Haibin Wen, Fanfu Wang et al.NeurIPS 2024 · 26 citations
- When Imbalance Meets Imbalance: Structure-driven Learning for Imbalanced Graph ClassificationWei Xu, Pengkun Wang, Zhe Zhao, Binwu Wang et al.WWW 2024 · 19 citations
- LogicTree: Improving Complex Reasoning of LLMs via Instantiated Multi-step Synthetic Logical DataZehao Wang, Lin F. Yang, Jie Wang, Kehan Wang et al.NeurIPS 2025 · 5 citations
- Fair Training with Zero InputsWenjie Pan, Jianqing Zhu, Huanqiang ZengAAAI 2025 · 2 citations
- Rethinking Long-tailed Dataset Distillation: A Uni-Level Framework with Unbiased Recovery and RelabelingXiao Cui, Yulei Qin, Xinyue Li, Wengang Zhou et al.AAAI 2026 · 1 citation
Builds on22
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan et al.ICLR 2020 · 1,496 citations
- Balanced Meta-Softmax for Long-Tailed Visual RecognitionJiawei Ren, Cunjun Yu, Shunan Sheng, Xiao Ma et al.NeurIPS 2020 · 861 citations
- Long-Tailed Classification by Keeping the Good and Removing the Bad Momentum Causal EffectKaihua Tang, Jianqiang Huang, Hanwang ZhangNeurIPS 2020 · 533 citations
- Long-tailed Recognition by Routing Diverse Distribution-Aware ExpertsXudong Wang, Long Lian, Zhongqi Miao, Ziwei Liu et al.ICLR 2021 · 481 citations
Related papers
- CUDA: Curriculum of Data Augmentation for Long-tailed RecognitionSumyeong Ahn, Jongwoo Ko, Se-Young YunICLR 2023 · 15 citations
- MetaSAug: Meta Semantic Augmentation for Long-Tailed Visual RecognitionShuang Li, Kaixiong Gong, Chi Harold Liu, Yulin Wang et al.CVPR 2021
- How Re-sampling Helps for Long-Tail Learning?Jiang-Xin Shi, Tong Wei, Yuke Xiang, Yufeng LiNeurIPS 2023 · 84 citations
- Enhancing Minority Classes by Mixing: An Adaptative Optimal Transport Approach for Long-tailed ClassificationJintong Gao, He Zhao, Zhuo Li, Dandan GuoNeurIPS 2023 · 64 citations
- Rethinking Class-Balanced Methods for Long-Tailed Visual Recognition From a Domain Adaptation PerspectiveMuhammad Abdullah Jamal, Matthew Brown, Ming-Hsuan Yang, Liqiang Wang et al.CVPR 2020
