Uniformly Distributed Category Prototype-Guided Vision-Language Framework for Long-Tail Recognition
Xiaoxuan He, Siming Fu, Xinpeng Ding, Yuchen Cao, Hualiang Wang
Abstract
Recently, large-scale pre-trained vision-language models have presented benefits for alleviating class imbalance in long-tailed recognition. However, the long-tailed data distribution can corrupt the representation space, where the distance between head and tail categories is much larger than the distance between two tail categories. This uneven feature space distribution causes the model to exhibit unclear and inseparable decision boundaries on the uniformly distributed test set, which lowers its performance. To address these challenges, we propose the uniformly category prototype-guided vision-language framework to effectively mitigate feature space bias caused by data imbalance. Especially, we generate a set of category prototypes uniformly distributed on a hypersphere. Category prototype-guided mechanism for image-text matching makes the features of different classes converge to these distinct and uniformly distributed category prototypes, which maintain a uniform distribution in the feature space, and improve class boundaries. Additionally, our proposed irrelevant text filtering and attribute enhancement module allows the model to ignore irrelevant noisy text and focus more on key attribute information, thereby enhancing the robustness of our framework. In the image recognition fine-tuning stage, to address the positive bias problem of the learnable classifier, we design the class feature prototype-guided classifier, which compensates for the performance of tail classes while maintaining the performance of head classes. Our method outperforms previous vision-language methods for long-tailed learning work by a large margin and achieves state-of-the-art performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 21348b0c-dd73-4614-9703-37d60a4cbc28Cited by top-tier papers3
- Long-Tail Learning with Foundation Model: Heavy Fine-Tuning HurtsJiang-Xin Shi, Tong Wei, Zhi Zhou, Jie-Jing Shao et al.ICML 2024 · 78 citations
- Bridge the Modality and Capability Gaps in Vision-Language Model SelectionChao Yi, Yuhang He, De-Chuan Zhan, Han-Jia YeNeurIPS 2024 · 32 citations
- RetFormer: Multimodal Retrieval for Enhancing Image RecognitionTianrui Yu, Xiubo Liang, Hongzhi WangCVPR 2026
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
Related papers
- HGLTR: Hierarchical Knowledge Injection for Calibrating Pre-trained Models in Long-Tail RecognitionJinpeng Zheng, Shao-Yuan Li, Gan Xu, Wenhai Wan et al.AAAI 2026
- Category-Prompt Refined Feature Learning for Long-Tailed Multi-Label Image ClassificationJiexuan Yan, Sheng Huang, Nankun Mu, Luwen Huangfu et al.ACM MM 2024 · 12 citations
- Adaptive Token Refinement in Long-Tailed Large Vision-Language Models Fine-TuningWenjun Miao, Mingda Li, Yanchao Hao, Zheng WeiICML 2026
- CLIP-Guided Federated Learning on Heterogeneity and Long-Tailed DataJiangming Shi, Shanshan Zheng, Xiangbo Yin, Yang Lu et al.AAAI 2024 · 38 citations
- Efficient and Long-Tailed Generalization for Pre-trained Vision-Language ModelJiang-Xin Shi, Chi Zhang, Tong Wei, Yufeng LiKDD 2024 · 3 citations
