Diff2I2P: Differentiable Image-to-Point Cloud Registration with Diffusion Prior
Juncheng Mu, Chengwei Ren, Weixiang Zhang, Liang Pan, Xiao-Ping Zhang, Yue Gao
Abstract
Learning cross-modal correspondences is essential for image-to-point cloud (I2P) registration. Existing methods achieve this mostly by utilizing metric learning to enforce feature alignment across modalities, disregarding the inherent modality gap between image and point data. Consequently, this paradigm struggles to ensure accurate crossmodal correspondences. To this end, inspired by the crossmodal generation success of recent large diffusion models, we propose Diff2 I2P, a fully Diff erentiable I2P registration framework, leveraging a novel and effective Diff usion prior for bridging the modality gap. Specifically, we propose a Control-Side Score Distillation (CSD) technique to distill knowledge from a depth-conditioned diffusion model to directly optimize the predicted transformation. However, the gradients on the transformation fail to backpropagate onto the cross-modal features due to the non-differentiability of correspondence retrieval and PnP solver. To this end, we further propose a Deformable Correspondence Tuning (DCT) module to estimate the correspondences in a differentiable way, followed by the transformation estimation using a differentiable PnP solver. With these two designs, the Diffusion model serves as a strong prior to guide the crossmodal feature learning of image and point cloud for forming robust correspondences, which significantly improves the registration. Extensive experimental results demonstrate that Diff2 I2P consistently outperforms SoTA I2P registration methods, achieving over 7% improvement in registration recall on the 7-Scenes benchmark. Code will be available at https://github.com/mujc2021/Diff2I2P.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e5105be7-6bdf-4302-9e38-18dc6621fef3Cited by top-tier papers4
- FS-I2P: A Hierarchical Focus–Sweep Registration Network with Dynamically Allocated DepthZhixin Cheng, Yujia Chen, Xujing Tao, Bohao Liao et al.ICML 2026 · 2 citations
- Rethinking 2D-3D Registration: A Novel Network for High-Value Zone Selection and Representation Consistency AlignmentZhixin Cheng, Bohao Liao, Jiacheng Deng, Xiaotian Yin et al.CVPR 2026 · 2 citations
- Hg-I2P: Bridging Modalities for Generalizable Image-to-Point-Cloud Registration via Heterogeneous GraphsPei An, Junfeng Ding, Jiaqi Yang, Yulong Wang et al.CVPR 2026 · 1 citation
- PlanaReLoc: Camera Relocalization in 3D Planar Primitives via Region-Based Structure MatchingHanqiao Ye, Yuzhou Liu, Yangdong Liu, Shuhan ShenCVPR 2026
Builds on30
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui et al.ICCV 2019 · 3,193 citations
- DreamGaussian: Generative Gaussian Splatting for Efficient 3D Content CreationJiaxiang Tang, Jiawei Ren, Hang Zhou, Ziwei Liu et al.ICLR 2024 · 955 citations
Related papers
- FreeReg: Image-to-Point Cloud Registration Leveraging Pretrained Diffusion Models and Monocular Depth EstimatorsHaiping Wang, Yuan Liu, Bing Wang, Yujing Sun et al.ICLR 2024 · 33 citations
- MinCD-PnP: Learning 2D-3D Correspondences with Approximate Blind PnPPei An, Jiaqi Yang, Muyao Peng, You Yang et al.ICCV 2025 · 5 citations
- RARE: Refine Any Registration of Pairwise Point Clouds via Zero-Shot LearningChengyu Zheng, Jin Huang, Honghua Chen, Mingqiang WeiICCV 2025 · 2 citations
- 2D3D-MATR: 2D-3D Matching Transformer for Detection-free Registration between Images and Point CloudsMinhao Li, Zheng Qin, Zhirui Gao, Renjiao Yi et al.ICCV 2023 · 30 citations
- Differentiable Registration of Images and LiDAR Point Clouds with VoxelPoint-to-Pixel MatchingJunsheng Zhou, Baorui Ma, Wenyuan Zhang, Yi Fang et al.NeurIPS 2023 · 62 citations
