Scene Grounding in the Wild
Tamir Cohen, Leo Segre, Shay Shomer Chai, Shai Avidan, Hadar Averbuch-Elor
摘要
Reconstructing accurate 3D models of large-scale real-world scenes from unstructured, in-the-wild imagery remains a core challenge in computer vision, especially when the input views have little or no overlap. In such cases, existing reconstruction pipelines often produce multiple disconnected partial reconstructions or erroneously merge non-overlapping regions into overlapping geometry. In this work, we propose a framework that grounds each partial reconstruction to a complete reference model of the scene, enabling globally consistent alignment even in the absence of visual overlap. We obtain reference models from dense, geospatially accurate pseudo-synthetic renderings derived from Google Earth Studio. These renderings provide full scene coverage but differ substantially in appearance from real-world photographs. Our key insight is that, despite this significant domain gap, both domains share the same underlying scene semantics. We represent the reference model using 3D Gaussian Splatting, augmenting each Gaussian with semantic features, and formulate alignment as an inverse feature-based optimization scheme that estimates a global 6DoF pose and scale while keeping the reference model fixed. Furthermore, we introduce the WikiEarth dataset, which registers existing partial 3D reconstructions with pseudo-synthetic reference models. We demonstrate that our approach consistently improves global alignment when initialized with various classical and learning-based pipelines, while mitigating failure modes of state-of-the-art end-to-end models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Deep Closest Point: Learning Representations for Point Cloud RegistrationYue Wang, Justin SolomonICCV 2019 · 被引用 1,026 次
- LightGlue: Local Feature Matching at Light SpeedPhilipp Lindenberger, Paul-Edouard Sarlin, Marc PollefeysICCV 2023 · 被引用 936 次
- Language-driven Semantic SegmentationBoyi Li, Kilian Q. Weinberger, Serge J. Belongie, Vladlen Koltun 等ICLR 2022 · 被引用 885 次
相关 Paper
- Cross-View Splatter: Feed-Forward View Synthesis with Georeferenced ImagesMatias Turkulainen, Akshay Krishnan, Filippo Aleotti, Mohamed Sayed 等CVPR 2026
- Cross-Instance Gaussian Splatting Registration via Geometry-Aware Feature-Guided AlignmentRoy Amoyal, Oren Freifeld, Chaim BaskinCVPR 2026
- SelfSplat: Pose-Free and 3D Prior-Free Generalizable 3D Gaussian SplattingGyeongjin Kang, Jisang Yoo, Jihyeon Park, Seungtae Nam 等CVPR 2025
- VGGS: VGGT-guided Gaussian Splatting for Efficient and Faithful Sparse-View Surface ReconstructionPeng Xiang, Liang Han, Hui Zhang, Yu-Shen Liu 等AAAI 2026 · 被引用 1 次
- PoseGaussian: 6D Pose Estimation for Unseen Objects via Sparse-View Object-Level 3D Gaussian SplattingWubin Shi, Shaoyan Gai, Feipeng DaCVPR 2026
