Scene Grounding in the Wild
Tamir Cohen, Leo Segre, Shay Shomer Chai, Shai Avidan, Hadar Averbuch-Elor
Abstract
Reconstructing accurate 3D models of large-scale real-world scenes from unstructured, in-the-wild imagery remains a core challenge in computer vision, especially when the input views have little or no overlap. In such cases, existing reconstruction pipelines often produce multiple disconnected partial reconstructions or erroneously merge non-overlapping regions into overlapping geometry. In this work, we propose a framework that grounds each partial reconstruction to a complete reference model of the scene, enabling globally consistent alignment even in the absence of visual overlap. We obtain reference models from dense, geospatially accurate pseudo-synthetic renderings derived from Google Earth Studio. These renderings provide full scene coverage but differ substantially in appearance from real-world photographs. Our key insight is that, despite this significant domain gap, both domains share the same underlying scene semantics. We represent the reference model using 3D Gaussian Splatting, augmenting each Gaussian with semantic features, and formulate alignment as an inverse feature-based optimization scheme that estimates a global 6DoF pose and scale while keeping the reference model fixed. Furthermore, we introduce the WikiEarth dataset, which registers existing partial 3D reconstructions with pseudo-synthetic reference models. We demonstrate that our approach consistently improves global alignment when initialized with various classical and learning-based pipelines, while mitigating failure modes of state-of-the-art end-to-end models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Deep Closest Point: Learning Representations for Point Cloud RegistrationYue Wang, Justin SolomonICCV 2019 · 1,026 citations
- LightGlue: Local Feature Matching at Light SpeedPhilipp Lindenberger, Paul-Edouard Sarlin, Marc PollefeysICCV 2023 · 936 citations
- Language-driven Semantic SegmentationBoyi Li, Kilian Q. Weinberger, Serge J. Belongie, Vladlen Koltun et al.ICLR 2022 · 885 citations
Related papers
- Cross-View Splatter: Feed-Forward View Synthesis with Georeferenced ImagesMatias Turkulainen, Akshay Krishnan, Filippo Aleotti, Mohamed Sayed et al.CVPR 2026
- Cross-Instance Gaussian Splatting Registration via Geometry-Aware Feature-Guided AlignmentRoy Amoyal, Oren Freifeld, Chaim BaskinCVPR 2026
- SelfSplat: Pose-Free and 3D Prior-Free Generalizable 3D Gaussian SplattingGyeongjin Kang, Jisang Yoo, Jihyeon Park, Seungtae Nam et al.CVPR 2025
- VGGS: VGGT-guided Gaussian Splatting for Efficient and Faithful Sparse-View Surface ReconstructionPeng Xiang, Liang Han, Hui Zhang, Yu-Shen Liu et al.AAAI 2026 · 1 citation
- PoseGaussian: 6D Pose Estimation for Unseen Objects via Sparse-View Object-Level 3D Gaussian SplattingWubin Shi, Shaoyan Gai, Feipeng DaCVPR 2026
