HGLTR: Hierarchical Knowledge Injection for Calibrating Pre-trained Models in Long-Tail Recognition
Jinpeng Zheng, Shao-Yuan Li, Gan Xu, Wenhai Wan, Zijian Tao, Songcan Chen, Kangkan Wang
Abstract
Long-tail recognition remains challenging for pre-trained foundation models like CLIP, which often suffer from performance degradation under imbalanced data. This stems not only from the overfitting/underfitting issues during fine-tuning but, more fundamentally, from the inherent bias inherited from the long-tail distribution of their massive pre-training datasets. To address this, we propose HGLTR (Hierarchy-Guided Long-Tail Recognition), a novel framework that calibrates pre-trained models by injecting objective class hierarchy knowledge. We argue that the semantic proximity defined by a hierarchy provides a robust, data-independent prior to counteract model bias. Our method is specifically designed for vision-language models' dual-modality architecture. At the feature level, we align image embeddings with a hierarchy-guided text similarity structure. At the classifier level, we employ a distillation loss to regularize predictions using soft labels derived from the hierarchy. This dual-level injection effectively transfers knowledge from head to tail classes. Experiments on ImageNet-LT, Places-LT, and iNaturalist 2018 demonstrate that HGLTR achieves state-of-the-art performance, particularly in tail-classes accuracy, highlighting the importance of leveraging structural priors to calibrate foundation models for real-world data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1760ca05-5e25-4dbd-9036-756dde8b0c8eBuilds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan et al.ICLR 2020 · 1,496 citations
- Long-tail learning via logit adjustmentAditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu Jain et al.ICLR 2021 · 937 citations
- Long-tailed Recognition by Routing Diverse Distribution-Aware ExpertsXudong Wang, Long Lian, Zhongqi Miao, Ziwei Liu et al.ICLR 2021 · 481 citations
Related papers
- Uniformly Distributed Category Prototype-Guided Vision-Language Framework for Long-Tail RecognitionXiaoxuan He, Siming Fu, Xinpeng Ding, Yuchen Cao et al.ACM MM 2023 · 6 citations
- Distributional Robustness Loss for Long-tail LearningDvir Samuel, Gal ChechikICCV 2021 · 128 citations
- Category-Prompt Refined Feature Learning for Long-Tailed Multi-Label Image ClassificationJiexuan Yan, Sheng Huang, Nankun Mu, Luwen Huangfu et al.ACM MM 2024 · 12 citations
- CLIP-Guided Federated Learning on Heterogeneity and Long-Tailed DataJiangming Shi, Shanshan Zheng, Xiangbo Yin, Yang Lu et al.AAAI 2024 · 38 citations
- CARE: Class-Adaptive Expert Consensus for Reliable Learning with Long-Tailed Noisy LabelsMengke Li, Haiquan Ling, Lihao Chen, Yang Lu et al.ICML 2026
