Open-Vocabulary Point-Cloud Object Detection without 3D Annotation
Yuheng Lu, Chenfeng Xu, Xiaobao Wei, Xiaodong Xie, Masayoshi Tomizuka, Kurt Keutzer, Shanghang Zhang
Abstract
The goal of open-vocabulary detection is to identify novel objects based on arbitrary textual descriptions. In this paper, we address open-vocabulary 3D point-cloud detection by a dividing-and-conquering strategy, which involves: 1) developing a point-cloud detector that can learn a general representation for localizing various objects, and 2) connecting textual and point-cloud representations to enable the detector to classify novel object categories based on text prompting. Specifically, we resort to rich image pre-trained models, by which the point-cloud detector learns localizing objects under the supervision of predicted 2D bounding boxes from 2D pre-trained detectors. Moreover, we propose a novel de-biased triplet cross-modal contrastive learning to connect the modalities of image, point-cloud and text, thereby enabling the point-cloud detector to benefit from vision-language pre-trained models, i.e., CLIP. The novel use of image and vision-language pretrained models for point-cloud detectors allows for openvocabulary 3D object detection without the need for 3D annotations. Experiments demonstrate that the proposed method improves at least 3.03 points and 7.47 points over a wide range of baselines on the ScanNet and SUN RGB-D datasets, respectively. Furthermore, we provide a comprehensive analysis to explain why our approach works. Code is available at https://github.com/lyhdet/OV-3DET
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f40d1734-e4a3-4e8d-b1e6-94952bea52c0Cited by top-tier papers35
- OpenShape: Scaling Up 3D Shape Representation Towards Open-World UnderstandingMinghua Liu, Ruoxi Shi, Kaiming Kuang, Yinhao Zhu et al.NeurIPS 2023 · 267 citations
- OpenGaussian: Towards Point-Level 3D Gaussian-based Open Vocabulary UnderstandingYanmin Wu, Jiarui Meng, Haijie Li, Chenming Wu et al.NeurIPS 2024 · 191 citations
- CoDA: Collaborative Novel Box Discovery and Cross-modal Alignment for Open-vocabulary 3D Object DetectionYang Cao, Yihan Zeng, Hang Xu, Dan XuNeurIPS 2023 · 69 citations
- Language Embedded 3D Gaussians for Open-Vocabulary Scene UnderstandingJin-Chuan Shi, Miao Wang, Hao-Bin Duan, Shao-Hua GuanCVPR 2024 · 57 citations
- Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask GuidancePhuc D. A. Nguyen, Tuan Duc Ngo, Evangelos Kalogerakis, Chuang Gan et al.CVPR 2024 · 45 citations
Builds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- An End-to-End Transformer Model for 3D Object DetectionIshan Misra, Rohit Girdhar, Armand JoulinICCV 2021 · 602 citations
Related papers
- PointCLIP: Point Cloud Understanding by CLIPRenrui Zhang, Ziyu Guo, Wei Zhang, Kunchang Li et al.CVPR 2022
- PointCLIP V2: Prompting CLIP and GPT for Powerful 3D Open-world LearningXiangyang Zhu, Renrui Zhang, Bowei He, Ziyu Guo et al.ICCV 2023 · 248 citations
- CLIP-Driven Open-Vocabulary 3D Scene Graph Generation via Cross-Modality Contrastive LearningLianggangxu Chen, Xuejiao Wang, Jiale Lu, Shaohui Lin et al.CVPR 2024
- Simple Image-Level Classification Improves Open-Vocabulary Object DetectionRuohuan Fang, Guansong Pang, Xiao BaiAAAI 2024 · 26 citations
- CLIP2Scene: Towards Label-efficient 3D Scene Understanding by CLIPRunnan Chen, Youquan Liu, Lingdong Kong, Xinge Zhu et al.CVPR 2023
