Vision-Language Pre-training with Object Contrastive Learning for 3D Scene Understanding
Taolin Zhang, Sunan He, Tao Dai, Zhi Wang, Bin Chen, Shu-Tao Xia
Abstract
In recent years, vision language pre-training frameworks have made significant progress in natural language processing and computer vision, achieving remarkable performance improvement on various downstream tasks. However, when extended to point cloud data, existing works mainly focus on building task-specific models, and fail to extract universal 3D vision-language embedding that generalize well. We carefully investigate three common tasks in semantic 3D scene understanding, and derive key insights into the development of a pre-training model. Motivated by these observations, we propose a vision-language pre-training framework 3DVLP (3D vision-language pre-training with object contrastive learning), which transfers flexibly on 3D vision-language downstream tasks. 3DVLP takes visual grounding as the proxy task and introduces Object-level IoU-guided Detection (OID) loss to obtain high-quality proposals in the scene. Moreover, we design Object-level Cross-Contrastive alignment (OCC) task and Object-level Self-Contrastive learning (OSC) task to align the objects with descriptions and distinguish different objects in the scene, respectively. Extensive experiments verify the excellent performance of 3DVLP on three 3D vision-language tasks, reflecting its superiority in semantic 3D scene understanding. CCS CONCEPTS • Computing methodologies → Computer vision; Natural language processing; Neural networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d95fa703-79b4-424b-8fcc-f79279799b7eCited by top-tier papers16
- Towards Compact 3D Representations via Point Feature Enhancement Masked AutoencodersYaohua Zha, Huizhen Ji, Jinmin Li, Rongsheng Li et al.AAAI 2024 · 66 citations
- Bridging the Gap between 2D and 3D Visual Question Answering: A Fusion Approach for 3D VQAWentao Mo, Yang LiuAAAI 2024 · 30 citations
- LCM: Locally Constrained Compact Point Cloud Model for Masked Point ModelingYaohua Zha, Naiqi Li, Yanzi Wang, Tao Dai et al.NeurIPS 2024 · 25 citations
- Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied NavigationZiyu Zhu, Xilin Wang, Yixuan Li, Zhuofan Zhang et al.ICCV 2025 · 11 citations
- LIBA: Language Instructed Multi-granularity Bridge Assistant for 3D Visual GroundingYuan Wang, Yali Li, Eastman Z. Y. Wu, Shengjin WangAAAI 2025 · 11 citations
Builds on26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
Related papers
- Point-GCC: Universal Self-supervised 3D Scene Pre-training via Geometry-Color ContrastGuofan Fan, Zekun Qi, Wenkai Shi, Kaisheng MaACM MM 2024 · 12 citations
- CLIP2Point: Transfer CLIP to Point Cloud Classification with Image-Depth Pre-TrainingTianyu Huang, Bowen Dong, Yunhan Yang, Xiaoshui Huang et al.ICCV 2023 · 220 citations
- FAC: 3D Representation Learning via Foreground Aware Feature ContrastKangcheng Liu, Aoran Xiao, Xiaoqin Zhang, Shijian Lu et al.CVPR 2023
- GLIPv2: Unifying Localization and Vision-Language UnderstandingHaotian Zhang, Pengchuan Zhang, Xiaowei Hu, Yen-Chun Chen et al.NeurIPS 2022 · 403 citations
- SpaceCLIP: A Vision-Language Pretraining Framework With Spatial Reconstruction On TextBo Zou, Chao Yang, Chengbin Quan, Youjian ZhaoACM MM 2023 · 1 citation
