SG2Loc: Sequential Visual Localization on 3D Scene Graphs
Nicole Damblon, Olga Vysotska, Federico Tombari, Marc Pollefeys, Daniel Barath
Abstract
Visual localization in complex indoor environments remains a critical challenge for robotics and AR applications. Sequential localization, where pose estimates are refined over time, is important for autonomous agents. However, traditional methods often require storing extensive image databases or point clouds, leading to significant overhead. This paper introduces a novel, lightweight approach to sequential visual localization using 3D scene graphs. Our method represents the environment with a compact scene graph, where nodes represent objects (with coarse meshes) and edges encode spatial relationships. For each image in the localization phase, we extract per-patch semantic features, predicting object identities. Localization is performed within a particle filter framework. Each particle, representing a camera pose, projects the coarse object meshes from the scene graph into the image, assigning object identities to patches based on visibility. The similarity of the per-patch features, in the input image, and object features from the scene graph determines the weight of a particle. Subsequent images are incorporated sequentially, refining the pose estimate. By leveraging a compact scene graph and efficient semantic matching, our method significantly reduces storage while maintaining performance on real-world datasets. The code will be available at https://github.com/DmblnNicole/sg2loc .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d8022c1c-a90b-4225-8e2e-9fc15dc7cc7bBuilds on12
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
- 3D Scene Graph: A Structure for Unified Semantics, 3D Space, and CameraIro Armeni, Zhi-Yang He, Amir Zamir, JunYoung Gwak et al.ICCV 2019 · 474 citations
- Learning With Average Precision: Training Image Retrieval With a Listwise LossJérôme Revaud, Jon Almazán, Rafael S. Rezende, César Roberto de SouzaICCV 2019 · 424 citations
- Deep Patch Visual OdometryZachary Teed, Lahav Lipson, Jia DengNeurIPS 2023 · 323 citations
- RIO: 3D Object Instance Re-Localization in Changing Indoor EnvironmentsJohanna Wald, Armen Avetisyan, Nassir Navab, Federico Tombari et al.ICCV 2019 · 233 citations
Related papers
- SceneSqueezer: Learning to Compress Scene for Camera RelocalizationLuwei Yang, Rakesh Shrestha, Wenbo Li, Shuaicheng Liu et al.CVPR 2022 · 29 citations
- SegLoc: Learning Segmentation-Based Representations for Privacy-Preserving Visual LocalizationMaxime Pietrantoni, Martin Humenberger, Torsten Sattler, Gabriela CsurkaCVPR 2023
- SAG-GNN: Semantic-Aware Guided GNN for Descriptor-Free 2D-3D MatchingShihua Zhang, Tianhao Xu, Zizhuo Li, Qing Ma et al.CVPR 2026
- VS-Net: Voting With Segmentation for Visual LocalizationZhaoyang Huang, Han Zhou, Yijin Li, Bangbang Yang et al.CVPR 2021
- Hierarchical 3D Scene Graphs Construction OutdoorsJon Nyffeler, Federico Tombari, Daniel BarathICCV 2025 · 1 citation
