Learning Imbalanced Data with Vision Transformers
Zhengzhuo Xu, Ruikang Liu, Shuo Yang, Zenghao Chai, Chun Yuan
Abstract
The real-world data tends to be heavily imbalanced and severely skew the data-driven deep neural networks, which makes Long-Tailed Recognition (LTR) a massive challenging task. Existing LTR methods seldom train Vision Transformers (ViTs) with Long-Tailed (LT) data, while the offthe-shelf pretrain weight of ViTs always leads to unfair comparisons. In this paper, we systematically investigate the ViTs' performance in LTR and propose LiVT to train ViTs from scratch only with LT data. With the observation that ViTs suffer more severe LTR problems, we conduct Masked Generative Pretraining (MGP) to learn generalized features. With ample and solid evidence, we show that MGP is more robust than supervised manners. Although Binary Cross Entropy (BCE) loss performs well with ViTs, it struggles on the LTR tasks. We further propose the balanced BCE to ameliorate it with strong theoretical groundings. Specially, we derive the unbiased extension of Sigmoid and compensate extra logit margins for deploying it. Our Bal-BCE contributes to the quick convergence of ViTs in just a few epochs. Extensive experiments demonstrate that with MGP and Bal-BCE, LiVT successfully trains ViTs well without any additional data and outperforms comparable state-of-the-art methods significantly, e.g., our ViT-B achieves 81.0% Top-1 accuracy in iNaturalist 2018 without bells and whistles. Code is available at https://github.com/XuZhengzhuo/LiVT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5afec789-eb98-4e0c-8d6b-dc66a96d63e7Cited by top-tier papers21
- REVE: A Foundation Model for EEG - Adapting to Any Setup with Large-Scale Pretraining on 25, 000 SubjectsYassine El Ouahidi, Jonathan Lys, Philipp Thölke, Nicolas Farrugia et al.NeurIPS 2025 · 106 citations
- Long-Tail Learning with Foundation Model: Heavy Fine-Tuning HurtsJiang-Xin Shi, Tong Wei, Zhi Zhou, Jie-Jing Shao et al.ICML 2024 · 78 citations
- Improving Visual Prompt Tuning by Gaussian Neighborhood Minimization for Long-Tailed Visual RecognitionMengke Li, Ye Liu, Yang Lu, Yiqun Zhang et al.NeurIPS 2024 · 27 citations
- NeurIPT: Foundation Model for Neural InterfacesZitao Fang, Chenxuan Li, Hongting Zhou, Shuyang Yu et al.NeurIPS 2025 · 16 citations
- DeiT-LT: Distillation Strikes Back for Vision Transformer Training on Long-Tailed DatasetsHarsh Rangwani, Pradipto Mondal, Mayank Mishra, Ashish Ramayee Asokan et al.CVPR 2024 · 12 citations
Builds on46
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
Related papers
- Distributional Robustness Loss for Long-tail LearningDvir Samuel, Gal ChechikICCV 2021 · 128 citations
- BCE3S: Binary Cross-Entropy Based Tripartite Synergistic Learning for Long-Tailed RecognitionWeijia Fan, Qiufu Li, Jiajun Wen, Xiaoyang PengAAAI 2026
- Balanced Contrastive Learning for Long-Tailed Visual RecognitionJianggang Zhu, Zheng Wang, Jingjing Chen, Yi-Ping Phoebe Chen et al.CVPR 2022 · 194 citations
- Towards Calibrated Model for Long-Tailed Visual Recognition from Prior PerspectiveZhengzhuo Xu, Zenghao Chai, Chun YuanNeurIPS 2021 · 77 citations
- HGLTR: Hierarchical Knowledge Injection for Calibrating Pre-trained Models in Long-Tail RecognitionJinpeng Zheng, Shao-Yuan Li, Gan Xu, Wenhai Wan et al.AAAI 2026
