Category-Prompt Refined Feature Learning for Long-Tailed Multi-Label Image Classification
Jiexuan Yan, Sheng Huang, Nankun Mu, Luwen Huangfu, Bo Liu
摘要
Real-world data consistently exhibits a long-tailed distribution, often spanning multiple categories. This complexity underscores the challenge of content comprehension, particularly in scenarios requiring Long-Tailed Multi-Label image Classification (LTMLC). In such contexts, imbalanced data distribution and multi-object recognition pose significant hurdles. To address this issue, we propose a novel and effective approach for LTMLC, termed Category-Prompt Refined Feature Learning (CPRFL), utilizing semantic correlations between different categories and decoupling category-specific visual representations for each category. Specifically, CPRFL initializes category-prompts from the pretrained CLIP's embeddings and decouples category-specific visual representations through interaction with visual features, thereby facilitating the establishment of semantic correlations between the head and tail classes. To mitigate the visual-semantic domain bias, we design a progressive Dual-Path Back-Propagation mechanism to refine the prompts by progressively incorporating context-related visual information into prompts. Simultaneously, the refinement process facilitates the progressive purification of the category-specific visual representations under the guidance of the refined prompts. Furthermore, taking into account the negative-positive sample imbalance, we adopt the Asymmetric Loss as our optimization objective to suppress negative samples across all classes and potentially enhance the head-to-tail recognition performance. We validate the effectiveness of our method on two LTMLC benchmarks and extensive experiments demonstrate the superiority of our work over baselines. The code is available at https://github.com/jiexuanyan/CPRFL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- MUSE: Harnessing Precise and Diverse Semantics for Few-Shot Whole Slide Image ClassificationJiahao Xu, Sheng Huang, Xin Zhang, Zhixiong Nan 等CVPR 2026 · 被引用 2 次
- Category-Specific Selective Feature Enhancement for Long-Tailed Multi-Label Image ClassificationRuiqi Du, Xu Tang, Xiangrong Zhang, Jingjing MaICCV 2025 · 被引用 1 次
- Rethinking BCE Loss for Multi-Label Image Recognition with Fine-TuningAo Zhou, Zhiwei Jiang, Zifeng Cheng, Cong Wang 等CVPR 2026
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan 等ICLR 2020 · 被引用 1,496 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
- Asymmetric Loss For Multi-Label ClassificationTal Ridnik, Emanuel Ben Baruch, Nadav Zamir, Asaf Noy 等ICCV 2021 · 被引用 778 次
相关 Paper
- Uniformly Distributed Category Prototype-Guided Vision-Language Framework for Long-Tail RecognitionXiaoxuan He, Siming Fu, Xinpeng Ding, Yuchen Cao 等ACM MM 2023 · 被引用 6 次
- HGLTR: Hierarchical Knowledge Injection for Calibrating Pre-trained Models in Long-Tail RecognitionJinpeng Zheng, Shao-Yuan Li, Gan Xu, Wenhai Wan 等AAAI 2026
- LPT: Long-tailed Prompt Tuning for Image ClassificationBowen Dong, Pan Zhou, Shuicheng Yan, Wangmeng ZuoICLR 2023 · 被引用 19 次
- Decoupled Contrastive Learning for Long-Tailed RecognitionShiyu Xuan, Shiliang ZhangAAAI 2024 · 被引用 29 次
- Parameter-Efficient Complementary Expert Learning for Long-Tailed Visual RecognitionLixiang Ru, Xin Guo, Lei Yu, Yingying Zhang 等ACM MM 2024 · 被引用 2 次
