Learning to Detect and Segment for Open Vocabulary Object Detection
Tao Wang
摘要
Open vocabulary object detection has been greatly advanced by the recent development of vision-language pretrained model, which helps recognize novel objects with only semantic categories. The prior works mainly focus on knowledge transferring to the object proposal classification and employ class-agnostic box and mask prediction. In this work, we propose CondHead, a principled dynamic network design to better generalize the box regression and mask segmentation for open vocabulary setting. The core idea is to conditionally parameterize the network heads on semantic embedding and thus the model is guided with class-specific knowledge to better detect novel categories. Specifically, CondHead is composed of two streams of network heads, the dynamically aggregated head and dynamically generated head. The former is instantiated with a set of static heads that are conditionally aggregated, these heads are optimized as experts and are expected to learn sophisticated prediction. The latter is instantiated with dynamically generated parameters and encodes general class-specific information. With such a conditional design, the detection model is bridged by the semantic embedding to offer strongly generalizable class-wise box and mask prediction. Our method brings significant improvement to the state-of-the-art open vocabulary object detection methods with very minor overhead, e.g., it surpasses a RegionClip model by 3.0 detection AP on novel categories, with only 1.1% more computation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Zero-shot Visual Relation Detection via Composite Visual Cues from Large Language ModelsLin Li, Jun Xiao, Guikun Chen, Jian Shao 等NeurIPS 2023 · 被引用 52 次
- Zero-Shot Aerial Object Detection with Visual Description RegularizationZhengqing Zang, Chenyu Lin, Chenwei Tang, Tao Wang 等AAAI 2024 · 被引用 22 次
- T-VSL: Text-Guided Visual Sound Source Localization in MixturesTanvir Mahmud, Yapeng Tian, Diana MarculescuCVPR 2024 · 被引用 8 次
- Scene-adaptive and Region-aware Multi-modal Prompt for Open Vocabulary Object DetectionXiaowei Zhao, Xianglong Liu, Duorui Wang, Yajun Gao 等CVPR 2024 · 被引用 8 次
- OVMR: Open-Vocabulary Recognition with Multi-Modal ReferencesZehong Ma, Shiliang Zhang, Longhui Wei, Qi TianCVPR 2024 · 被引用 6 次
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 被引用 1,274 次
- Objects365: A Large-Scale, High-Quality Dataset for Object DetectionShuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng 等ICCV 2019 · 被引用 1,018 次
相关 Paper
- CORA: Adapting CLIP for Open-Vocabulary Detection with Region Prompting and Anchor Pre-MatchingXiaoshi Wu, Feng Zhu, Rui Zhao, Hongsheng LiCVPR 2023
- Side Adapter Network for Open-Vocabulary Semantic SegmentationMengde Xu, Zheng Zhang, Fangyun Wei, Han Hu 等CVPR 2023
- Distilling DETR with Visual-Linguistic Knowledge for Open-Vocabulary Object DetectionLiangqi Li, Jiaxu Miao, Dahu Shi, Wenming Tan 等ICCV 2023 · 被引用 35 次
- CAKE: Category Aware Knowledge Extraction for Open-Vocabulary Object DetectionShiyuan Ma, Donglin Qian, Kai Ye, Shengchuan ZhangAAAI 2025 · 被引用 8 次
- Simple Image-Level Classification Improves Open-Vocabulary Object DetectionRuohuan Fang, Guansong Pang, Xiao BaiAAAI 2024 · 被引用 26 次
