Language-Driven Multi-Label Zero-Shot Learning with Semantic Granularity
Shouwen Wang, Qian Wan, Junbin Gao, Zhigang Zeng
摘要
Recent methods learn class-unified prompt contexts by image data to adapt CLIP to zero-shot multi-label image classification, which achieves impressive performance. However, simply tuning prompts is insufficient to deal with novel classes across different semantic granularity levels. This limitation arises due to the sparse semantic detail in prompt class names and the hierarchical granularity competition among class names caused by CLIP's contrastive loss. We propose a language-driven zero-shot multi-label learning framework to bridge associations among classes across multiple granularity levels through class name reconstruction. To achieve this, we first leverage a language model to generate structured text descriptions for each class, which explicitly capture (1) visual attributes, (2) hierarchical relationships, and (3) co-occurrence scenes. With the enriched descriptions, we then learn class names by extracting and aligning semantic relationships and features from them in the CLIP's shared image-text embedding space. Furthermore, we consider that similar text descriptions among different classes may introduce ambiguities. We mitigate these ambiguities by imposing a pair-based loss on learnable class names to enhance their distinctiveness. During inference, we aggregate semantic predictions from multiple image snippets to reinforce the identification of classes across different granularity levels. Comprehensive experiments demonstrate that our method surpasses state-of-theart methods in multi-label zero-shot learning and effectively handles novel classes across different granularity levels.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Learning Semantic-Specific Graph Representation for Multi-Label Image RecognitionTianshui Chen, Muxin Xu, Xiaolu Hui, Hefeng Wu 等ICCV 2019 · 被引用 347 次
- DualCoOp: Fast Adaptation to Multi-Label Recognition with Limited AnnotationsXimeng Sun, Ping Hu, Kate SaenkoNeurIPS 2022 · 被引用 199 次
- Multi-Label Classification with Label Graph SuperimposingYa Wang, Dongliang He, Fu Li, Xiang Long 等AAAI 2020 · 被引用 192 次
相关 Paper
- Rethinking the Effect of Uninformative Class Name in Prompt LearningFengmao Lv, Changru Nie, Jianyang Zhang, Guowu Yang 等ACM MM 2024 · 被引用 1 次
- Interpreting and Analysing CLIP's Zero-Shot Image Classification via Mutual KnowledgeFawaz Sammani, Nikos DeligiannisNeurIPS 2024 · 被引用 15 次
- ChatGPT-Powered Hierarchical Comparisons for Image ClassificationZhiyuan Ren, Yiyang Su, Xiaoming LiuNeurIPS 2023 · 被引用 54 次
- Follow-Up Differential Descriptions: Language Models Resolve Ambiguities for Image ClassificationReza Esfandiarpoor, Stephen H. BachICLR 2024 · 被引用 18 次
- ViLT-CLIP: Video and Language Tuning CLIP with Multimodal Prompt Learning and Scenario-Guided OptimizationHao Wang, Fang Liu, Licheng Jiao, Jiahao Wang 等AAAI 2024 · 被引用 54 次
