Improved Visual-Semantic Alignment for Zero-Shot Object Detection
Shafin Rahman, Salman H. Khan, Nick Barnes
Abstract
Zero-shot object detection is an emerging research topic that aims to recognize and localize previously 'unseen' objects. This setting gives rise to several unique challenges, e.g., highly imbalanced positive vs. negative instance ratio, proper alignment between visual and semantic concepts and the ambiguity between background and unseen classes. Here, we propose an end-to-end deep learning framework underpinned by a novel loss function that handles class-imbalance and seeks to properly align the visual and semantic cues for improved zero-shot learning. We call our objective the 'Polarity loss' because it explicitly maximizes the gap between positive and negative predictions. Such a margin maximizing formulation is not only important for visual-semantic alignment but it also resolves the ambiguity between background and unseen objects. Further, the semantic representations of objects are noisy, thus complicating the alignment between visual and semantic domains. To this end, we perform metric learning using a 'Semantic vocabulary' of related concepts that refines the noisy semantic embeddings and establishes a better synergy between visual and semantic domains. Our approach is inspired by the embodiment theories in cognitive science, that claim human semantic understanding to be grounded in past experiences (seen objects), related linguistic concepts (word vocabulary) and the visual perception (seen/unseen object images). Our extensive results on MS-COCO and Pascal VOC datasets show significant improvements over state of the art. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers38
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li et al.CVPR 2022 · 481 citations
- Bridging the Gap between Object and Image-level Representations for Open-Vocabulary DetectionHanoona Abdul Rasheed, Muhammad Maaz, Muhammad Uzair Khattak, Salman H. Khan et al.NeurIPS 2022 · 215 citations
- Fine-Grained Semantically Aligned Vision-Language Pre-TrainingJuncheng Li, Xin He, Longhui Wei, Long Qian et al.NeurIPS 2022 · 111 citations
- Open-Vocabulary Instance Segmentation via Robust Cross-Modal Pseudo-LabelingDat Huynh, Jason Kuen, Zhe Lin, Jiuxiang Gu et al.CVPR 2022 · 78 citations
Builds on1
Related papers
- Zero-Shot Object Detection by Semantics-Aware DETR with Adaptive Contrastive LossHuan Liu, Lu Zhang, Jihong Guan, Shuigeng ZhouACM MM 2023 · 6 citations
- Meta-ZSDETR: Zero-shot DETR with Meta-learningLu Zhang, Chenbo Zhang, Jiajia Zhao, Jihong Guan et al.ICCV 2023 · 10 citations
- Semantic-Promoted Debiasing and Background Disambiguation for Zero-Shot Instance SegmentationShuting He, Henghui Ding, Wei JiangCVPR 2023
- Robust Region Feature Synthesizer for Zero-Shot Object DetectionPeiliang Huang, Junwei Han, De Cheng, Dingwen ZhangCVPR 2022 · 50 citations
- Recognizing Unseen Objects via Multimodal Intensive Knowledge Graph PropagationLikang Wu, Zhi Li, Hongke Zhao, Zhefeng Wang et al.KDD 2023 · 4 citations
