Sculpting Holistic 3D Representation in Contrastive Language-Image-3D Pre-Training
Yipeng Gao, Zeyu Wang, Wei-Shi Zheng, Cihang Xie, Yuyin Zhou
Abstract
Contrastive learning has emerged as a promising paradigm for 3D open-world understanding, i.e., aligning point cloud representation to image and text embedding space individually. In this paper, we introduce Mix-Con3D, a simple yet effective method aiming to sculpt holistic 3D representation in contrastive language-image-3D pre-training. In contrast to point cloud only, we develop the 3D object-level representation from complementary perspectives, e.g., multi-view rendered images with the point cloud. Then, MixCon3D performs language-3D contrastive learning, comprehensively depicting real-world 3D objects and bolstering text alignment. Additionally, we pioneer the first thorough investigation of various training recipes for the 3D contrastive learning paradigm, building a solid baseline with improved performance. Extensive experiments conducted on three representative benchmarks reveal that our method significantly improves over the baseline, surpassing the previous state-of-the-art performance on the challenging 1,156-category Objaverse-LVIS dataset by 5.7%. The versatility of MixCon3D is showcased in applications such as text-to-3D retrieval and point cloud captioning, further evidencing its efficacy in diverse scenarios. The code is available at https://github.com/UCSC-VLAA/MixCon3D .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Panoptic Captioning: An Equivalence Bridge for Image and TextKun-Yu Lin, Hongjun Wang, Weining Ren, Kai HanNeurIPS 2025 · 7 citations
- More Text, Less Point: Towards 3D Data-Efficient Point-Language UnderstandingYuan Tang, Xu Han, Xianzhi Li, Qiao Yu et al.AAAI 2025 · 6 citations
- Point Cloud as a Foreign Language for Multi-modal Large Language ModelSneha Paul, Zachary Patterson, Nizar BouguilaCVPR 2026 · 2 citations
- TIGaussian: Disentangle Gaussians for Spatial-Awared Text-Image-3D AlignmentJiarun Liu, Qifeng Chen, Yiru Zhao, Minghua Liu et al.ICLR 2026 · 1 citation
- LAM: Language Articulated Object ModelersYipeng Gao, Yunhao Ge, Peilin Cai, Daniel Seita et al.CVPR 2026
Builds on48
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui et al.ICCV 2019 · 3,193 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
Related papers
- OpenShape: Scaling Up 3D Shape Representation Towards Open-World UnderstandingMinghua Liu, Ruoxi Shi, Kaiming Kuang, Yinhao Zhu et al.NeurIPS 2023 · 267 citations
- MM-Mixing: Multi-Modal Mixing Alignment for 3D UnderstandingJiaze Wang, Yi Wang, Ziyu Guo, Renrui Zhang et al.AAAI 2025 · 1 citation
- RegionPLC: Regional Point-Language Contrastive Learning for Open-World 3D Scene UnderstandingJihan Yang, Runyu Ding, Weipeng Deng, Zhe Wang et al.CVPR 2024 · 49 citations
- CLIP2Point: Transfer CLIP to Point Cloud Classification with Image-Depth Pre-TrainingTianyu Huang, Bowen Dong, Yunhan Yang, Xiaoshui Huang et al.ICCV 2023 · 220 citations
- Vision-Language Pre-training with Object Contrastive Learning for 3D Scene UnderstandingTaolin Zhang, Sunan He, Tao Dai, Zhi Wang et al.AAAI 2024 · 42 citations
