HTTM: Head-wise Temporal Token Merging for Faster VGGT
Weitian Wang, Lukas Meiner, Shubham Rai, Cecilia De la Parra, Akash Kumar
摘要
The Visual Geometry Grounded Transformer (VGGT) marks a significant leap forward in 3D scene reconstruction, as it is the first model that directly infers all key 3D attributes (camera poses, depths, and dense geometry) jointly in one pass. However, this joint inference mechanism requires global attention layers that perform all-to-all attention computation on tokens from all views. For reconstruction of large scenes with long-sequence inputs, this causes a significant latency bottleneck. In this paper, we propose head-wise temporal merging (HTTM), a training-free 3D token merging method for accelerating VGGT. Existing merging techniques merge tokens uniformly across different attention heads, resulting in identical tokens in the layers'output, which hinders the model's representational ability. HTTM tackles this problem by merging tokens in multi-head granularity, which preserves the uniqueness of feature tokens after head concatenation. Additionally, this enables HTTM to leverage the spatial locality and temporal correspondence observed at the head level to achieve higher merging ratios with lower merging costs compared to existing methods. Thus, HTTM achieves up to acceleration over the original VGGT with negligible performance drops in a GPU-based inference.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
相关 Paper
- LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token MergingZhijian Shu, Cheng Lin, Tao Xie, Wei Yin 等CVPR 2026 · 被引用 17 次
- FastVGGT: Fast Visual Geometry TransformerYou Shen, Zhipeng Zhang, Yansong Qu, Xiawu Zheng 等ICLR 2026 · 被引用 73 次
- Co-Me: Confidence Guided Token Merging for Visual Geometric TransformersYutian Chen, Yuheng Qiu, Ruogu Li, Jay Patrikar 等CVPR 2026 · 被引用 9 次
- QVGGT: Post-Training Quantized Visual Geometry Grounded TransformerZhizhen Pan, Hesong Wang, Huan WangCVPR 2026 · 被引用 2 次
- FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor AttentionZipeng Wang, Dan XuCVPR 2026 · 被引用 14 次
