Semantic Diversity Learning for Zero-Shot Multi-label Classification
Avi Ben-Cohen, Nadav Zamir, Emanuel Ben Baruch, Itamar Friedman, Lihi Zelnik-Manor
Abstract
Training a neural network model for recognizing multiple labels associated with an image, including identifying unseen labels, is challenging, especially for images that portray numerous semantically diverse labels. As challenging as this task is, it is an essential task to tackle since it represents many real-world cases, such as image retrieval of natural images. We argue that using a single embedding vector to represent an image, as commonly practiced, is not sufficient to rank both relevant seen and unseen labels accurately. This study introduces an end-to-end model training for multi-label zero-shot learning that supports the semantic diversity of the images and labels. We propose to use an embedding matrix having principal embedding vectors trained using a tailored loss function. In addition, during training, we suggest up-weighting in the loss function image samples presenting higher semantic diversity to encourage the diversity of the embedding matrix. Extensive experiments show that our proposed method improves the zero-shot model’s quality in tag-based image retrieval achieving SoTA results on several common datasets (NUS-Wide, COCO, Open Images).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- DualCoOp: Fast Adaptation to Multi-Label Recognition with Limited AnnotationsXimeng Sun, Ping Hu, Kate SaenkoNeurIPS 2022 · 199 citations
- Open-Vocabulary Multi-Label Classification via Multi-Modal Knowledge TransferSunan He, Taian Guo, Tao Dai, Ruizhi Qiao et al.AAAI 2023 · 76 citations
- Discriminative Region-based Multi-Label Zero-Shot LearningSanath Narayan, Akshita Gupta, Salman H. Khan, Fahad Shahbaz Khan et al.ICCV 2021 · 62 citations
- TagCLIP: A Local-to-Global Framework to Enhance Open-Vocabulary Multi-Label Classification of CLIP without TrainingYuqi Lin, Minghao Chen, Kaipeng Zhang, Hengjia Li et al.AAAI 2024 · 39 citations
- Simple Image-Level Classification Improves Open-Vocabulary Object DetectionRuohuan Fang, Guansong Pang, Xiao BaiAAAI 2024 · 26 citations
Builds on5
- Asymmetric Loss For Multi-Label ClassificationTal Ridnik, Emanuel Ben Baruch, Nadav Zamir, Asaf Noy et al.ICCV 2021 · 778 citations
- Cross-Modality Attention with Semantic Graph Embedding for Multi-Label ClassificationRenchun You, Zhiyao Guo, Lei Cui, Xiang Long et al.AAAI 2020 · 221 citations
- Multi-Label Classification with Label Graph SuperimposingYa Wang, Dongliang He, Fu Li, Xiang Long et al.AAAI 2020 · 192 citations
- Transductive Learning for Zero-Shot Object DetectionShafin Rahman, Salman H. Khan, Nick BarnesICCV 2019 · 82 citations
- A Shared Multi-Attention Framework for Multi-Label Zero-Shot LearningDat Huynh, Ehsan ElhamifarCVPR 2020
Related papers
- Epsilon: Exploring Comprehensive Visual-Semantic Projection for Multi-Label Zero-Shot LearningZiming Liu, Jingcai Guo, Song Guo, Xiaocheng LuAAAI 2025 · 6 citations
- Adaptive and Generative Zero-Shot LearningYu-Ying Chou, Hsuan-Tien Lin, Tyng-Luh LiuICLR 2021 · 25 citations
- Inferring Prototypes for Multi-Label Few-Shot Image Classification with Word Vector Guided AttentionKun Yan, Chenbin Zhang, Jun Hou, Ping Wang et al.AAAI 2022 · 27 citations
- Semantic-Promoted Debiasing and Background Disambiguation for Zero-Shot Instance SegmentationShuting He, Henghui Ding, Wei JiangCVPR 2023
- Robust Region Feature Synthesizer for Zero-Shot Object DetectionPeiliang Huang, Junwei Han, De Cheng, Dingwen ZhangCVPR 2022 · 50 citations
