EP2P-Loc: End-to-End 3D Point to 2D Pixel Localization for Large-Scale Visual Localization
Minjung Kim, Junseo Koo, Gunhee Kim
摘要
Visual localization is the task of estimating a 6-DoF camera pose of a query image within a provided 3D reference map. Thanks to recent advances in various 3D sensors, 3D point clouds are becoming a more accurate and affordable option for building the reference map, but research to match the points of 3D point clouds with pixels in 2D images for visual localization remains challenging. Existing approaches that jointly learn 2D-3D feature matching suffer from low inliers due to representational differences between the two modalities, and the methods that bypass this problem into classification have an issue of poor refinement. In this work, we propose EP2P-Loc, a novel large-scale visual localization method that mitigates such appearance discrepancy and enables end-to-end training for pose estimation. To increase the number of inliers, we propose a simple algorithm to remove invisible 3D points in the image, and find all 2D-3D correspondences without keypoint detection. To reduce memory usage and search complexity, we take a coarse-to-fine approach where we extract patch-level features from 2D images, then perform 2D patch classification on each 3D point, and obtain the exact corresponding 2D pixel coordinates through positional encoding. Finally, for the first time in this task, we employ a differentiable PnP for end-to-end training. In the experiments on newly curated large-scale indoor and outdoor benchmarks based on 2D-3D-S and KITTI, we show that our method achieves the state-of-the-art performance compared to existing visual localization and image-to-point cloud registration methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- FreeReg: Image-to-Point Cloud Registration Leveraging Pretrained Diffusion Models and Monocular Depth EstimatorsHaiping Wang, Yuan Liu, Bing Wang, Yujing Sun 等ICLR 2024 · 被引用 33 次
- MinCD-PnP: Learning 2D-3D Correspondences with Approximate Blind PnPPei An, Jiaqi Yang, Muyao Peng, You Yang 等ICCV 2025 · 被引用 5 次
- ConDo: Continual Domain Expansion for Absolute Pose RegressionZijun Li, Zhipeng Cai, Bochun Yang, Xuelun Shen 等AAAI 2025 · 被引用 2 次
- Trafficloc: Localizing Traffic Surveillance Cameras in 3D ScenesYan Xia, Yunxiang Lu, Rui Song, Oussema Dhaouadi 等ICCV 2025 · 被引用 2 次
- NormalLoc: Visual Localization on Textureless 3D Models using Surface NormalsJiro Abe, Gaku Nakano, Kazumine OguraICCV 2025 · 被引用 2 次
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 被引用 2,416 次
- USIP: Unsupervised Stable Interest Point Detection From 3D Point CloudsJiaxin Li, Gim Hee LeeICCV 2019 · 被引用 206 次
- Unified Contrastive Learning in Image-Text-Label SpaceJianwei Yang, Chunyuan Li, Pengchuan Zhang, Bin Xiao 等CVPR 2022 · 被引用 182 次
相关 Paper
- P2-Net: Joint Description and Detection of Local Features for Pixel and Point MatchingBing Wang, Changhao Chen, Zhaopeng Cui, Jie Qin 等ICCV 2021 · 被引用 75 次
- RayI2P: Learning Rays for Image-to-Point Cloud RegistrationXinjun Li, Wenfei Yang, Zhixin Cheng, Jiacheng Deng 等ICLR 2026
- Back to the Feature: Learning Robust Camera Localization From Pixels To PosePaul-Edouard Sarlin, Ajaykumar Unagar, Måns Larsson, Hugo Germain 等CVPR 2021
- DenserNet: Weakly Supervised Visual Localization Using Multi-Scale Feature AggregationDongfang Liu, Yiming Cui, Liqi Yan, Christos Mousas 等AAAI 2021 · 被引用 149 次
- Learning Multi-View Aggregation In the Wild for Large-Scale 3D Semantic SegmentationDamien Robert, Bruno Vallet, Loïc LandrieuCVPR 2022 · 被引用 84 次
