Selfi: Self-improving Reconstruction Engine via 3D Geometric Feature Alignment
Youming Deng, Songyou Peng, Junyi Zhang, Kathryn Heal, Tiancheng Sun, John Flynn, Steve Marschner, Lucy Chai
Abstract
Novel View Synthesis (NVS) has traditionally relied on models with explicit 3D inductive biases combined with known camera parameters from Structure-from-Motion (SfM) beforehand. Recent vision foundation models like VGGT take an orthogonal approach -- 3D knowledge is gained implicitly through training data and loss objectives, enabling feed-forward prediction of both camera parameters and 3D representations directly from a set of uncalibrated images. While flexible, VGGT features lack explicit multi-view geometric consistency, and we find that improving such 3D feature consistency benefits both NVS and pose estimation tasks. We introduce Selfi, a self-improving 3D reconstruction pipeline via feature alignment, transforming a VGGT backbone into a high-fidelity 3D reconstruction engine by leveraging its own outputs as pseudo-ground-truth. Specifically, we train a lightweight feature adapter using a reprojection-based consistency loss, which distills VGGT outputs into a new geometrically-aligned feature space that captures spatial proximity in 3D. This enables state-of-the-art performance in both NVS and camera pose estimation, demonstrating that feature alignment is a highly beneficial step for downstream 3D reasoning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 94f3560d-7db1-4277-806b-9ef59d59ef2fBuilds on43
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan et al.CVPR 2022 · 1,603 citations
Related papers
- VGGT: Visual Geometry Grounded TransformerJianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi et al.CVPR 2025
- GGPT: Geometry-Grounded Point TransformerYutong Chen, Yiming Wang, Xucong Zhang, Sergey Prokudin et al.CVPR 2026 · 2 citations
- 3DPoV: Improving 3D understanding via Patch Ordering on VideosIoana Simion, Mohammadreza Salehi, Shashanka Venkataramanan, Cees Snoek et al.ICML 2026
- VGGT-ΩJianyuan Wang, Minghao Chen, Shangzhan Zhang, Nikita Karaev et al.CVPR 2026
- Gen3R: 3D Scene Generation Meets Feed-Forward ReconstructionJiaxin Huang, Yuanbo Yang, Bangbang Yang, Lin Ma et al.CVPR 2026 · 24 citations
