Learning to Detect Scene Landmarks for Camera Localization
Tien Do, Ondrej Miksik, Joseph DeGol, Hyun Soo Park, Sudipta N. Sinha
Abstract
In this supplementary document, we present additional quantitative results that could not be included in the main paper. We show extensive qualitative results from our method on the INDOOR-6 dataset and also include a supplemental video. Finally, we discuss some failure cases. Quantitative Results In this section, we show the storage efficiency of our method (NBE+SLD) compared to a retrieval and matchingbased method (HLoc [3]) (Section 1.1). Next, we further compare accuracy between our method and multiple baselines through a recall plot that uses a range of thresholds (Section 1.2). Finally, we report bearing errors for predicted landmarks on INDOOR-6 and 7-SCENES [6] datasets (Section 1.3). Storage comparison NBE+SLD requires constant storage. Figure 1 reports the storage requirements for HLoc and our method for each scene in the INDOOR-6 dataset. Our method requires 0.135 GB of storage for the SLD and NBE networks' parameters that are constant for all the scenes. This is significantly smaller than HLoc that requires 1.5GB and 1.2GB on the two larger scenes -scene1 and scene5, respectively. HLoc stores SuperPoint [1] and SuperGlue [4] networks' parameters and SuperPoint's features and VLAD [2] image descriptors for all the database images. The storage for features grows linearly with the number of database images, and that can dominate total storage on large scenes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3741e315-17ca-4d0e-b368-e70577077387Cited by top-tier papers13
- SplatLoc: 3D Gaussian Splatting-based Visual Localization for Augmented RealityHongjia Zhai, Xiyu Zhang, Boming Zhao, Hai Li et al.IEEE VR 2025 · 29 citations
- ACE-G: Improving Generalization of Scene Coordinate Regression Through Query Pre-TrainingLeonard Bruns, Axel Barroso-Laguna, Tommaso Cavallari, Áron Monszpart et al.ICCV 2025 · 5 citations
- A Scene is Worth a Thousand Features: Feed-Forward Camera Localization from a Collection of Image FeaturesAxel Barroso-Laguna, Tommaso Cavallari, Victor Prisacariu, Eric BrachmannICLR 2026 · 3 citations
- Scene Coordinate Reconstruction PriorsWenjing Bian, Axel Barroso-Laguna, Tommaso Cavallari, Victor Adrian Prisacariu et al.ICCV 2025 · 2 citations
- Semantic-Guided Camera Ray Regression for Visual LocalizationYesheng Zhang, Xu ZhaoICCV 2025 · 2 citations
Builds on1
Related papers
- Relational Context Learning for Human-Object Interaction DetectionSanghyun Kim, Deunsol Jung, Minsu ChoCVPR 2023
- DSFNet: Dual Space Fusion Network for Occlusion-Robust 3D Dense Face AlignmentHeyuan Li, Bo Wang, Yu Cheng, Mohan S. Kankanhalli et al.CVPR 2023
- GaVS: 3D-Grounded Video Stabilization via Temporally-Consistent Local Reconstruction and RenderingZinuo You, Stamatios Georgoulis, Anpei Chen, Siyu Tang et al.SIGGRAPH 2025 · 3 citations
- VideoLoc: Video-based Indoor Localization with Text InformationShusheng Li, Wenbo HeINFOCOM 2021 · 4 citations
- Liberated-Gs: 3D Gaussian Splatting Independent From Sfm Point CloudsWeihong Pan, Xiaoyu Zhang, Hongjia Zhai, Xiaojun Xiang et al.ICCV 2025 · 6 citations
