CLIP2: Contrastive Language-Image-Point Pretraining from Real-World Point Cloud Data
Yihan Zeng, Chenhan Jiang, Jiageng Mao, Jianhua Han, Chaoqiang Ye, Qingqiu Huang, Dit-Yan Yeung, Zhen Yang, Xiaodan Liang, Hang Xu
2023Year
36Top-tier citations
Abstract
tion. Experimental results on both indoor and outdoor scenarios show that our learned 3D representation has great transfer ability in downstream tasks, including zero-shot and few-shot 3D recognition, which boosts the state-of-theart methods by large margins. Furthermore, we provide analyses of the capability of different representations in real scenarios and present the optional ensemble scheme.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a00d0dca-d41f-4998-8cce-398e2e44cfbcCited by top-tier papers36
- LRM: Large Reconstruction Model for Single Image to 3DYicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi et al.ICLR 2024 · 813 citations
- Uni3D: Exploring Unified 3D Representation at ScaleJunsheng Zhou, Jinsheng Wang, Baorui Ma, Yu-Shen Liu et al.ICLR 2024 · 207 citations
- HS-FPN: High Frequency and Spatial Perception FPN for Tiny Object DetectionZican Shi, Jing Hu, Jie Ren, Hengkang Ye et al.AAAI 2025 · 85 citations
- CoDA: Collaborative Novel Box Discovery and Cross-modal Alignment for Open-vocabulary 3D Object DetectionYang Cao, Yihan Zeng, Hang Xu, Dan XuNeurIPS 2023 · 69 citations
- RegionPLC: Regional Point-Language Contrastive Learning for Open-World 3D Scene UnderstandingJihan Yang, Runyu Ding, Weipeng Deng, Zhe Wang et al.CVPR 2024 · 49 citations
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- Revisiting Point Cloud Classification: A New Benchmark Dataset and Classification Model on Real-World DataMikaela Angelina Uy, Quang-Hieu Pham, Binh-Son Hua, Duc Thanh Nguyen et al.ICCV 2019 · 1,003 citations
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point ModelingXumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang et al.CVPR 2022 · 684 citations
Related papers
- Generative Densification: Learning to Densify Gaussians for High-Fidelity Generalizable 3D ReconstructionSeungtae Nam, Xiangyu Sun, Gyeongjin Kang, Younggeun Lee et al.CVPR 2025
- Detect Anything 3D in the WildHanxue Zhang, Haoran Jiang, Qingsong Yao, Yanan Sun et al.ICCV 2025 · 6 citations
- Disentangling 3D Prototypical Networks for Few-Shot Concept LearningMihir Prabhudesai, Shamit Lal, Darshan Patil, Hsiao-Yu Tung et al.ICLR 2021 · 22 citations
- On the Transfer of Object-Centric Representation LearningAniket Rajiv Didolkar, Andrii Zadaianchuk, Anirudh Goyal, Michael Curtis Mozer et al.ICLR 2025
- Few-Shot Incremental 3D Object Detection in Dynamic Indoor EnvironmentsYun Zhu, Jianjun Qian, Jian Yang, Jin Xie et al.CVPR 2026 · 2 citations
