Improved Visual-Semantic Alignment for Zero-Shot Object Detection
Shafin Rahman, Salman H. Khan, Nick Barnes
摘要
Zero-shot object detection is an emerging research topic that aims to recognize and localize previously 'unseen' objects. This setting gives rise to several unique challenges, e.g., highly imbalanced positive vs. negative instance ratio, proper alignment between visual and semantic concepts and the ambiguity between background and unseen classes. Here, we propose an end-to-end deep learning framework underpinned by a novel loss function that handles class-imbalance and seeks to properly align the visual and semantic cues for improved zero-shot learning. We call our objective the 'Polarity loss' because it explicitly maximizes the gap between positive and negative predictions. Such a margin maximizing formulation is not only important for visual-semantic alignment but it also resolves the ambiguity between background and unseen objects. Further, the semantic representations of objects are noisy, thus complicating the alignment between visual and semantic domains. To this end, we perform metric learning using a 'Semantic vocabulary' of related concepts that refines the noisy semantic embeddings and establishes a better synergy between visual and semantic domains. Our approach is inspired by the embodiment theories in cognitive science, that claim human semantic understanding to be grounded in past experiences (seen objects), related linguistic concepts (word vocabulary) and the visual perception (seen/unseen object images). Our extensive results on MS-COCO and Pascal VOC datasets show significant improvements over state of the art. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper38
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 被引用 1,274 次
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li 等CVPR 2022 · 被引用 481 次
- Bridging the Gap between Object and Image-level Representations for Open-Vocabulary DetectionHanoona Abdul Rasheed, Muhammad Maaz, Muhammad Uzair Khattak, Salman H. Khan 等NeurIPS 2022 · 被引用 215 次
- Fine-Grained Semantically Aligned Vision-Language Pre-TrainingJuncheng Li, Xin He, Longhui Wei, Long Qian 等NeurIPS 2022 · 被引用 111 次
- Open-Vocabulary Instance Segmentation via Robust Cross-Modal Pseudo-LabelingDat Huynh, Jason Kuen, Zhe Lin, Jiuxiang Gu 等CVPR 2022 · 被引用 78 次
它引用的顶会 Paper1
相关 Paper
- Zero-Shot Object Detection by Semantics-Aware DETR with Adaptive Contrastive LossHuan Liu, Lu Zhang, Jihong Guan, Shuigeng ZhouACM MM 2023 · 被引用 6 次
- Meta-ZSDETR: Zero-shot DETR with Meta-learningLu Zhang, Chenbo Zhang, Jiajia Zhao, Jihong Guan 等ICCV 2023 · 被引用 10 次
- Semantic-Promoted Debiasing and Background Disambiguation for Zero-Shot Instance SegmentationShuting He, Henghui Ding, Wei JiangCVPR 2023
- Robust Region Feature Synthesizer for Zero-Shot Object DetectionPeiliang Huang, Junwei Han, De Cheng, Dingwen ZhangCVPR 2022 · 被引用 50 次
- Recognizing Unseen Objects via Multimodal Intensive Knowledge Graph PropagationLikang Wu, Zhi Li, Hongke Zhao, Zhefeng Wang 等KDD 2023 · 被引用 4 次
