Language-Driven Multi-Label Zero-Shot Learning with Semantic Granularity
Shouwen Wang, Qian Wan, Junbin Gao, Zhigang Zeng
Abstract
Recent methods learn class-unified prompt contexts by image data to adapt CLIP to zero-shot multi-label image classification, which achieves impressive performance. However, simply tuning prompts is insufficient to deal with novel classes across different semantic granularity levels. This limitation arises due to the sparse semantic detail in prompt class names and the hierarchical granularity competition among class names caused by CLIP's contrastive loss. We propose a language-driven zero-shot multi-label learning framework to bridge associations among classes across multiple granularity levels through class name reconstruction. To achieve this, we first leverage a language model to generate structured text descriptions for each class, which explicitly capture (1) visual attributes, (2) hierarchical relationships, and (3) co-occurrence scenes. With the enriched descriptions, we then learn class names by extracting and aligning semantic relationships and features from them in the CLIP's shared image-text embedding space. Furthermore, we consider that similar text descriptions among different classes may introduce ambiguities. We mitigate these ambiguities by imposing a pair-based loss on learnable class names to enhance their distinctiveness. During inference, we aggregate semantic predictions from multiple image snippets to reinforce the identification of classes across different granularity levels. Comprehensive experiments demonstrate that our method surpasses state-of-theart methods in multi-label zero-shot learning and effectively handles novel classes across different granularity levels.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2d1d0ab6-dbc1-4d94-82e0-9322c9529f31Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Learning Semantic-Specific Graph Representation for Multi-Label Image RecognitionTianshui Chen, Muxin Xu, Xiaolu Hui, Hefeng Wu et al.ICCV 2019 · 347 citations
- DualCoOp: Fast Adaptation to Multi-Label Recognition with Limited AnnotationsXimeng Sun, Ping Hu, Kate SaenkoNeurIPS 2022 · 199 citations
- Multi-Label Classification with Label Graph SuperimposingYa Wang, Dongliang He, Fu Li, Xiang Long et al.AAAI 2020 · 192 citations
Related papers
- Rethinking the Effect of Uninformative Class Name in Prompt LearningFengmao Lv, Changru Nie, Jianyang Zhang, Guowu Yang et al.ACM MM 2024 · 1 citation
- Interpreting and Analysing CLIP's Zero-Shot Image Classification via Mutual KnowledgeFawaz Sammani, Nikos DeligiannisNeurIPS 2024 · 15 citations
- ChatGPT-Powered Hierarchical Comparisons for Image ClassificationZhiyuan Ren, Yiyang Su, Xiaoming LiuNeurIPS 2023 · 54 citations
- Follow-Up Differential Descriptions: Language Models Resolve Ambiguities for Image ClassificationReza Esfandiarpoor, Stephen H. BachICLR 2024 · 18 citations
- ViLT-CLIP: Video and Language Tuning CLIP with Multimodal Prompt Learning and Scenario-Guided OptimizationHao Wang, Fang Liu, Licheng Jiao, Jiahao Wang et al.AAAI 2024 · 54 citations
