FastVGGT: Fast Visual Geometry Transformer
You Shen, Zhipeng Zhang, Yansong Qu, Xiawu Zheng, Jiayi Ji, Shengchuan Zhang, Liujuan Cao
摘要
Scaling visual geometry transformers for long image sequences poses a significant computational and memory challenge. In this work, we diagnose this issue in the state-of-the-art model VGGT, and trace the primary bottleneck to its Global Attention layer. Our analysis reveals a ``token collapse'' phenomenon, where many tokens attend to nearly identical regions, resulting in redundant computation and inefficiency. Motivated by this finding, we propose FastVGGT, a training-free framework that strategically prunes these redundant tokens. Instead of uniform merging, FastVGGT employs a tailored, three-part token partitioning strategy. It preserves initial-frame tokens as a stable global reference, retains salient tokens to maintain fine details, and utilizes region-based random sampling to ensure spatially balanced coverage. Extensive experiments on multiple 3D geometry benchmarks validate our approach's effectiveness. Notably, on sequences of 1000 images, FastVGGT achieves a 4 speedup over the original VGGT while simultaneously mitigating error accumulation, demonstrating its efficiency and robustness for long-sequence scenarios. For further details, please visit our project page: https://mystorm16.github.io/fastvggt/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Fast-FoundationStereo: Real-Time Zero-Shot Stereo MatchingBowen Wen, Shaurya Dewan, Stan BirchfieldCVPR 2026 · 被引用 36 次
- Scal3R: Scalable Test-Time Training for Large-Scale 3D ReconstructionTao Xie, Peishan Yang, Yudong Jin, Yingfeng Cai 等CVPR 2026 · 被引用 26 次
- ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time TrainingHaian Jin, Rundi Wu, Tianyuan Zhang, Ruiqi Gao 等CVPR 2026 · 被引用 23 次
- LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token MergingZhijian Shu, Cheng Lin, Tao Xie, Wei Yin 等CVPR 2026 · 被引用 17 次
- AVGGT: Rethinking Global Attention for Accelerating VGGTXianbing Sun, Zhikai Zhu, Zhengyu Lou, Bo Yang 等CVPR 2026 · 被引用 16 次
它引用的顶会 Paper18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii 等CVPR 2024 · 被引用 302 次
- Not All Tokens Are Equal: Human-centric Visual Analysis via Token Clustering TransformerWang Zeng, Sheng Jin, Wentao Liu, Chen Qian 等CVPR 2022 · 被引用 132 次
相关 Paper
- IncVGGT: Incremental VGGT for Memory-Bounded Long-Range 3D ReconstructionKeyu Fang, Changchun Zhou, Yuzhe Fu, Hai Li 等ICLR 2026
- HTTM: Head-wise Temporal Token Merging for Faster VGGTWeitian Wang, Lukas Meiner, Shubham Rai, Cecilia De la Parra 等CVPR 2026 · 被引用 7 次
- FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor AttentionZipeng Wang, Dan XuCVPR 2026 · 被引用 14 次
- Co-Me: Confidence Guided Token Merging for Visual Geometric TransformersYutian Chen, Yuheng Qiu, Ruogu Li, Jay Patrikar 等CVPR 2026 · 被引用 9 次
- QVGGT: Post-Training Quantized Visual Geometry Grounded TransformerZhizhen Pan, Hesong Wang, Huan WangCVPR 2026 · 被引用 2 次
