CLIP-Driven Open-Vocabulary 3D Scene Graph Generation via Cross-Modality Contrastive Learning
Lianggangxu Chen, Xuejiao Wang, Jiale Lu, Shaohui Lin, Changbo Wang, Gaoqi He
Abstract
3D Scene Graph Generation (3DSGG) aims to classify objects and their predicates within 3D point cloud scenes. However, current 3DSGG methods struggle with two main challenges. 1) The dependency on labor-intensive groundtruth annotations. 2) Closed-set classes training hampers the recognition of novel objects and predicates. Addressing these issues, our idea is to extract cross-modality features by CLIP from text and image data naturally related to 3D point clouds. Cross-modality features are used to train a robust 3D scene graph (3DSG) feature extractor. Specifically, we propose a novel Cross-Modality Contrastive Learning 3DSGG (CCL-3DSGG) method. Firstly, to align the text with 3DSG, the text is parsed into word level that are consistent with the 3DSG annotation. To enhance robustness during the alignment, adjectives are exchanged for different objects as negative samples. Then, to align the image with 3DSG, the camera view is treated as a positive sample and other views as negatives. Lastly, the recognition of novel object and predicate classes is achieved by calculating the cosine similarity between prompts and 3DSG features. Our rigorous experiments confirm the superior open-vocabulary capability and applicability of CCL-3DSGG in real-world contexts. Text Encoder Image Encoder CLIP I3D Loss T3D Loss 3DSG Feature Extractor Point Cloud (a) Difference in training (b) Difference in inference 3DSG Feature Extractor 3DSGG Model Collision Likelihood: ?
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2499481a-9c43-4f37-8d22-0aae194f5e9cCited by top-tier papers11
- ImaginateAR: AI-Assisted In-Situ Authoring in Augmented RealityJaewook Lee, Filippo Aleotti, Diego Mazala, Guillermo Garcia-Hernando et al.UIST 2025 · 15 citations
- Voxify3D: Pixel Art Meets Volumetric RenderingYi-Chuan Huang, Jiewen Chan, Hao-Jen Chien, Yu-Lun LiuCVPR 2026 · 3 citations
- Universal Scene Graph GenerationShengqiong Wu, Hao Fei, Tat-Seng ChuaCVPR 2025
- RelationField: Relate Anything in Radiance FieldsSebastian Koch, Johanna Wald, Mirco Colosi, Narunas Vaskevicius et al.CVPR 2025
- Learning 4D Panoptic Scene Graph Generation from Rich 2D Visual SceneShengqiong Wu, Hao Fei, Jingkang Yang, Xiangtai Li et al.CVPR 2025
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- 3D Scene Graph: A Structure for Unified Semantics, 3D Space, and CameraIro Armeni, Zhi-Yang He, Amir Zamir, JunYoung Gwak et al.ICCV 2019 · 474 citations
- NuScenes-QA: A Multi-Modal Visual Question Answering Benchmark for Autonomous Driving ScenarioTianwen Qian, Jingjing Chen, Linhai Zhuo, Yang Jiao et al.AAAI 2024 · 314 citations
- PointCLIP V2: Prompting CLIP and GPT for Powerful 3D Open-world LearningXiangyang Zhu, Renrui Zhang, Bowei He, Ziyu Guo et al.ICCV 2023 · 248 citations
Related papers
- Open-Vocabulary Point-Cloud Object Detection without 3D AnnotationYuheng Lu, Chenfeng Xu, Xiaobao Wei, Xiaodong Xie et al.CVPR 2023
- Multi-view Invariance Learning for 3D Scene Graph Pre-training via Collaborative Cross-Modal RegularizationYucheng Huang, Luping Ji, Ruijie Xiao, Jiayuan SunAAAI 2026
- All in One: Visual-Description-Guided Unified Point Cloud SegmentationZongyan Han, Mohamed El Amine Boudjoghra, Jiahua Dong, Jinhong Wang et al.ICCV 2025 · 1 citation
- CLIP2Scene: Towards Label-efficient 3D Scene Understanding by CLIPRunnan Chen, Youquan Liu, Lingdong Kong, Xinge Zhu et al.CVPR 2023
- Object-Centric Representation Learning for Enhanced 3D Semantic Scene Graph PredictionKunHo Heo, Gihyun Kim, SuYeon Kim, MyeongAh ChoNeurIPS 2025 · 4 citations
