Zero-shot RGB-D Point Cloud Registration with Pre-trained Large Vision Model
Haobo Jiang, Jin Xie, Jian Yang, Liang Yu, Jianmin Zheng
Abstract
This paper introduces ZeroMatch, a novel zero-shot RGB-D point cloud registration framework, aimed at achieving robust 3D matching on unseen data without any task-specific training. Our core idea is to utilize the powerful zeroshot image representation of Stable Diffusion, achieved through extensive pre-training on large-scale data, to enhance point-cloud geometric descriptors for robust matching. Specifically, we combine the handcrafted geometric descriptor FPFH with Stable-Diffusion features to create point descriptors that are both locally and contextually aware, enabling reliable RGB-D registration with zero-shot capability. This approach is based on our observation that Stable-Diffusion features effectively encode discriminative global contextual cues, naturally alleviating the feature ambiguity that FPFH often encounters in scenes with repetitive patterns or low overlap. To further enhance cross-view consistency of Stable-Diffusion features for improved matching, we propose a coupled-image input mode that concatenates the source and target images into a single input, replacing the original single-image mode. This design achieves both inter-image and prompt-to-image consistency attentions, facilitating robust cross-view feature interaction and alignment. Finally, we leverage feature nearest neighbors to construct putative correspondences for hypothesize-andverify transformation estimation. Extensive experiments on 3DMatch, ScanNet, and ScanLoNet verify the excellent zero-shot matching ability of our method. [Code]
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75f57035-fd40-4c91-8bf6-09ca4bb30f8eCited by top-tier papers6
- FUSER: Feed-Forward Multiview 3D Registration Transformer and SE(3)^N Diffusion RefinementHaobo Jiang, Jin Xie, Jian Yang, Liang Yu et al.CVPR 2026 · 5 citations
- RARE: Refine Any Registration of Pairwise Point Clouds via Zero-Shot LearningChengyu Zheng, Jin Huang, Honghua Chen, Mingqiang WeiICCV 2025 · 2 citations
- C-GenReg: Training-Free 3D Point Cloud Registration by Multi-View-Consistent Geometry-to-Image Generation with Probabilistic Modalities FusionYuval Haitman, Amit Efraim, Joseph M. FrancosCVPR 2026
- GM-R^2: Generative Matching Learning for Unsupervised Geometric Representation and RegistrationHaobo Jiang, Liang Yu, Jianmin ZhengCVPR 2026
- RGGT: A Generative-Prior-Guided Transformer for Unified Rigid and Non-Rigid Point Cloud RegistrationChengyu Zheng, Songlin Yang, Jin Huang, Honghua Chen et al.ICML 2026
Builds on24
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Fully Convolutional Geometric FeaturesChristopher B. Choy, Jaesik Park, Vladlen KoltunICCV 2019 · 807 citations
- Geometric Transformer for Fast and Robust Point Cloud RegistrationZheng Qin, Hao Yu, Changjian Wang, Yulan Guo et al.CVPR 2022 · 436 citations
- REGTR: End-to-end Point Cloud Correspondences with TransformersZi Jian Yew, Gim Hee LeeCVPR 2022 · 242 citations
Related papers
- FreeReg: Image-to-Point Cloud Registration Leveraging Pretrained Diffusion Models and Monocular Depth EstimatorsHaiping Wang, Yuan Liu, Bing Wang, Yujing Sun et al.ICLR 2024 · 33 citations
- Fuse2Match: Training-Free Fusion of Flow, Diffusion, and Contrastive Models for Zero-Shot Semantic MatchingJing Zuo, Jiaqi Wang, Yonggang Qi, Yi-Zhe SongNeurIPS 2025
- Harnessing Text-to-Image Diffusion Models for Point Cloud Self-Supervised LearningYiyang Chen, Shanshan Zhao, Lunhao Duan, Changxing Ding et al.ICCV 2025
- Buffer-X: Towards Zero-Shot Point Cloud Registration in Diverse ScenesMinkyun Seo, Hyungtae Lim, Kanghee Lee, Luca Carlone et al.ICCV 2025 · 9 citations
- Zero-Shot Point Cloud Segmentation by Semantic-Visual Aware SynthesisYuwei Yang, Munawar Hayat, Zhao Jin, Hongyuan Zhu et al.ICCV 2023 · 11 citations
