AVGGT: Rethinking Global Attention for Accelerating VGGT
Xianbing Sun, Zhikai Zhu, Zhengyu Lou, Bo Yang, Jinyang Tang, Liqing Zhang, He Wang, Jianfu Zhang
Abstract
Since DUSt3R, models such as VGGT and have shown strong multi-view 3D performance, but their heavy reliance on global self-attention results in high computational cost. Existing sparse-attention variants offer partial speedups, yet lack a systematic analysis of how global attention contributes to multi-view reasoning. In this paper, we first conduct an in-depth investigation of the global attention modules in VGGT and to better understand their roles. Our analysis reveals a clear division of roles in the alternating global-frame architecture: early global layers do not form meaningful correspondences, middle layers perform cross-view alignment, and last layers provide only minor refinements. Guided by these findings, we propose a training-free two-step acceleration scheme: (1) converting early global layers into frame attention, and (2) subsampling global attention by subsampling K/V over patch tokens with diagonal preservation and a mean-fill component.We instantiate this strategy on VGGT and and evaluate across standard pose and point-map benchmarks. Our method achieves up to - speedup in inference time while matching or slightly improving the accuracy of the original models, and remains robust even in extremely dense multi-view settings where prior sparse-attention baselines fail.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on15
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 769 citations
- π3: Permutation-Equivariant Visual Geometry LearningYifan Wang, Jianjun Zhou, Haoyi Zhu, Wenzheng Chang et al.ICLR 2026 · 318 citations
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii et al.CVPR 2024 · 302 citations
Related papers
- Block-Sparse Global Attention for Efficient Multi-View Geometry TransformersChung-Shien Brian Wang, Christian Schmidt, Jens Piekenbrinck, Bastian LeibeCVPR 2026 · 22 citations
- FastVGGT: Fast Visual Geometry TransformerYou Shen, Zhipeng Zhang, Yansong Qu, Xiawu Zheng et al.ICLR 2026 · 73 citations
- MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 SecondsZhenggang Tang, Yuchen Fan, Dilin Wang, Hongyu Xu et al.CVPR 2025
- HTTM: Head-wise Temporal Token Merging for Faster VGGTWeitian Wang, Lukas Meiner, Shubham Rai, Cecilia De la Parra et al.CVPR 2026 · 7 citations
- DVGT: Driving Visual Geometry TransformerSicheng Zuo, Zixun Xie, Wenzhao Zheng, Shaoqing Xu et al.CVPR 2026 · 23 citations
