Block-Sparse Global Attention for Efficient Multi-View Geometry Transformers
Chung-Shien Brian Wang, Christian Schmidt, Jens Piekenbrinck, Bastian Leibe
Abstract
Efficient and accurate feed-forward multi-view reconstruction has long been an important task in computer vision. Recent transformer-based models like VGGT, and MapAnything have demonstrated remarkable performance with relatively simple architectures. However, their scalability is fundamentally constrained by the quadratic complexity of global attention, which imposes a significant runtime bottleneck when processing large image sets. In this work, we empirically analyze the global attention matrix of these models and observe that the probability mass concentrates on a small subset of patch-patch interactions corresponding to cross-view geometric correspondences. Building on this insight and inspired by recent advances in large language models, we propose a training-free, block-sparse replacement for dense global attention, implemented with highly optimized kernels. Our method accelerates inference by more than while maintaining comparable task performance. Evaluations on a comprehensive suite of multi-view benchmarks demonstrate that our approach seamlessly integrates into existing global attention-based architectures such as VGGT, , and MapAnything, while substantially improving scalability to large image collections.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 24c24bda-d738-45c5-8eb2-a79fea638671Cited by top-tier papers8
- ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time TrainingHaian Jin, Rundi Wu, Tianyuan Zhang, Ruiqi Gao et al.CVPR 2026 · 23 citations
- AMB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with BackendHengyi Wang, Lourdes AgapitoCVPR 2026 · 17 citations
- AVGGT: Rethinking Global Attention for Accelerating VGGTXianbing Sun, Zhikai Zhu, Zhengyu Lou, Bo Yang et al.CVPR 2026 · 16 citations
- MVGGT: Multimodal Visual Geometry Grounded Transformer for Multiview 3D Referring Expression SegmentationChangli Wu, Haodong Wang, Jiayi Ji, Yutian Yao et al.CVPR 2026 · 8 citations
- KV-Tracker: Real-Time Pose Tracking with TransformersMarwan Taher, Ignacio Alzugaray, Kirill Mazur, Xin Kong et al.CVPR 2026 · 5 citations
Builds on29
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
Related papers
- FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor AttentionZipeng Wang, Dan XuCVPR 2026 · 14 citations
- FastVGGT: Fast Visual Geometry TransformerYou Shen, Zhipeng Zhang, Yansong Qu, Xiawu Zheng et al.ICLR 2026 · 73 citations
- HTTM: Head-wise Temporal Token Merging for Faster VGGTWeitian Wang, Lukas Meiner, Shubham Rai, Cecilia De la Parra et al.CVPR 2026 · 7 citations
- Hyena Hierarchy: Towards Larger Convolutional Language ModelsMichael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y. Fu et al.ICML 2023 · 481 citations
- GTA: A Geometry-Aware Attention Mechanism for Multi-View TransformersTakeru Miyato, Bernhard Jaeger, Max Welling, Andreas GeigerICLR 2024 · 51 citations
