Emergent Outlier View Rejection in Visual Geometry Grounded Transformers
Jisang Han, Sunghwan Hong, Jaewoo Jung, Wooseok Jang, Honggyu An, Qianqian Wang, Seungryong Kim, Chen Feng
Abstract
Reliable 3D reconstruction from in-the-wild image collections is often hindered by noisy images—irrelevant inputs with little or no view overlap with others. While traditional Structure-from-Motion pipelines handle such cases through geometric verification and outlier rejection, feed-forward 3D reconstruction models lack these explicit mechanisms, leading to degraded performance under in-the-wild conditions. In this paper, we discover that the existing feed-forward reconstruction model, e.g., VGGT, despite lacking explicit outlier-rejection mechanisms or noise-aware training, can inherently distinguish distractor images. Through an in-depth analysis under varying proportions of synthetic distractors, we identify a specific layer that naturally exhibits outlier-suppressing behavior. Further probing reveals that this layer encodes discriminative internal representations that enable an effective noise-filtering capability, which we simply leverage to perform outlier-view rejection in feed-forward 3D reconstruction without any additional fine-tuning or supervision. Extensive experiments on both controlled and in-the-wild datasets demonstrate that this implicit filtering mechanism is consistent and generalizes well across diverse scenarios. Code will be released.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2ab95757-453b-4119-9029-6bcea6911bd7Cited by top-tier papers1
Ask how each one uses itBuilds on26
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
- π3: Permutation-Equivariant Visual Geometry LearningYifan Wang, Jianjun Zhou, Haoyi Zhu, Wenzheng Chang et al.ICLR 2026 · 318 citations
- DeepV2D: Video to Depth with Differentiable Structure from MotionZachary Teed, Jia DengICLR 2020 · 314 citations
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii et al.CVPR 2024 · 302 citations
- CATs: Cost Aggregation Transformers for Visual CorrespondenceSeokju Cho, Sunghwan Hong, Sangryul Jeon, Yunsung Lee et al.NeurIPS 2021 · 133 citations
Related papers
- GGPT: Geometry-Grounded Point TransformerYutong Chen, Yiming Wang, Xucong Zhang, Sergey Prokudin et al.CVPR 2026 · 2 citations
- VGGT: Visual Geometry Grounded TransformerJianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi et al.CVPR 2025
- Selfi: Self-improving Reconstruction Engine via 3D Geometric Feature AlignmentYouming Deng, Songyou Peng, Junyi Zhang, Kathryn Heal et al.CVPR 2026 · 4 citations
- VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object DetectionYang Cao, Feize Wu, Dave Chen, Yingji Zhong et al.CVPR 2026 · 6 citations
- V-DPM: 4D Video Reconstruction with Dynamic Point MapsEdgar Sucar, Eldar Insafutdinov, Zihang Lai, Andrea VedaldiCVPR 2026 · 29 citations
