VGG-T: Offline Feed-Forward 3D Reconstruction at Scale
Sven Elflein, Ruilong Li, Sérgio Agostinho, Zan Gojcic, Laura Leal-Taixé, Qunjie Zhou, Aljosa Osep
2026年份
摘要
Forward time (s) OOM VGGT FastVGGT SparseVGGT Ours (b) Num. images vs. inference time.
Figure 1. Reconstructing Rome landmarks with 1-minute time budget.
We present VGG-T 3 , an offline feed-forward 3D reconstruction method that scales linearly w.r.t. input views (Fig. 1b). As a result, we can reconstruct large scenes from a large number of unposed input views, such as landmarks from tourist-sourced images, in less than a minute via single forward pass (Fig. 1a).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper55
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 被引用 1,407 次
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller 等ICML 2020 · 被引用 1,220 次
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski 等ICML 2023 · 被引用 848 次
- Common Objects in 3D: Large-Scale Learning and Evaluation of Real-life 3D Category ReconstructionJeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone 等ICCV 2021 · 被引用 686 次
相关 Paper
- VGGT: Visual Geometry Grounded TransformerJianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi 等CVPR 2025
- ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time TrainingHaian Jin, Rundi Wu, Tianyuan Zhang, Ruiqi Gao 等CVPR 2026 · 被引用 23 次
- Styl3R: Instant 3D Stylized Reconstruction for Arbitrary Scenes and StylesPeng Wang, Xiang Liu, Peidong LiuNeurIPS 2025 · 被引用 8 次
- Generalizable Sparse-View 3D Reconstruction from Unconstrained ImagesVinayak Gupta, Chih-Hao Lin, Shenlong Wang, Anand Bhattad 等CVPR 2026 · 被引用 1 次
- FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor AttentionZipeng Wang, Dan XuCVPR 2026 · 被引用 14 次
