SAG-GNN: Semantic-Aware Guided GNN for Descriptor-Free 2D-3D Matching
Shihua Zhang, Tianhao Xu, Zizhuo Li, Qing Ma, Jiayi Ma
Abstract
Image-to-point cloud matching (2D-3D matching) establishes accurate correspondences between image keypoints and 3D points for 6-DoF camera pose estimation. Existing methods either suffer from poor generalization due to scene-specific coordinate regression requiring per-scene retraining, or incur high storage and maintenance costs from descriptor-based matching that relies on large descriptor sets. Consequently, descriptor-free approaches have gained attention by avoiding heavy storage while improving generalizability; however, most rely only on low-level geometric cues, which limits performance. Leveraging the benefits of semantics in providing context, resolving ambiguities, and enhancing robustness in challenging scenes, we propose the Semantic-Aware Guided Graph Neural Network (SAG-GNN), integrating high-level semantics into descriptor-free 2D-3D matching. Specifically, we design a compact semantic extraction scheme encoding each 3D point as a lowdimensional semantic probability distribution, offering effective guidance with minimal storage. A bidirectionallyaligned fusion block merges geometric features with semantic context for more unified and consistent representations. Additionally, semantic priors guide the 2D-3D information exchange within the interaction framework from a high-level semantic perspective. Extensive indoor and outdoor experiments validate that SAG-GNN achieves stateof-the-art results in descriptor-free 2D-3D matching and visual localization, with low storage and strong generalization. Code is available at https://github.com/ tinxu0203/SAG-GNN .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ee4a992b-9b05-4198-be60-a7be470c5dccBuilds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- SegNeXt: Rethinking Convolutional Attention Design for Semantic SegmentationMeng-Hao Guo, Cheng-Ze Lu, Qibin Hou, Zhengning Liu et al.NeurIPS 2022 · 1,385 citations
- Learning Multi-Scene Absolute Pose Regression with TransformersYoli Shavit, Ron Ferens, Yosi KellerICCV 2021 · 163 citations
- Progressive Correspondence Pruning by Consensus LearningChen Zhao, Yixiao Ge, Feng Zhu, Rui Zhao et al.ICCV 2021 · 101 citations
- GLACE: Global Local Accelerated Coordinate EncodingFangjinhua Wang, Xudong Jiang, Silvano Galliani, Christoph Vogel et al.CVPR 2024 · 24 citations
Related papers
- DGC-GNN: Leveraging Geometry and Color Cues for Visual Descriptor-Free 2D-3D MatchingShuzhe Wang, Juho Kannala, Daniel BarathCVPR 2024 · 7 citations
- Learning Camera Localization via Dense Scene MatchingShitao Tang, Chengzhou Tang, Rui Huang, Siyu Zhu et al.CVPR 2021
- SGAD: Semantic and Geometric-Aware Descriptor for Local Feature MatchingXiangzeng Liu, Chi Wang, Guanglu Shi, Xiaodong Zhang et al.ICCV 2025 · 2 citations
- RayI2P: Learning Rays for Image-to-Point Cloud RegistrationXinjun Li, Wenfei Yang, Zhixin Cheng, Jiacheng Deng et al.ICLR 2026
- SG-PGM: Partial Graph Matching Network with Semantic Geometric Fusion for 3D Scene Graph Alignment and its Downstream TasksYaxu Xie, Alain Pagani, Didier StrickerCVPR 2024 · 5 citations
