HTTM: Head-wise Temporal Token Merging for Faster VGGT
Weitian Wang, Lukas Meiner, Shubham Rai, Cecilia De la Parra, Akash Kumar
Abstract
The Visual Geometry Grounded Transformer (VGGT) marks a significant leap forward in 3D scene reconstruction, as it is the first model that directly infers all key 3D attributes (camera poses, depths, and dense geometry) jointly in one pass. However, this joint inference mechanism requires global attention layers that perform all-to-all attention computation on tokens from all views. For reconstruction of large scenes with long-sequence inputs, this causes a significant latency bottleneck. In this paper, we propose head-wise temporal merging (HTTM), a training-free 3D token merging method for accelerating VGGT. Existing merging techniques merge tokens uniformly across different attention heads, resulting in identical tokens in the layers'output, which hinders the model's representational ability. HTTM tackles this problem by merging tokens in multi-head granularity, which preserves the uniqueness of feature tokens after head concatenation. Additionally, this enables HTTM to leverage the spatial locality and temporal correspondence observed at the head level to achieve higher merging ratios with lower merging costs compared to existing methods. Thus, HTTM achieves up to acceleration over the original VGGT with negligible performance drops in a GPU-based inference.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d0728eb2-d816-47a9-a337-723f36261aeeBuilds on17
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
Related papers
- LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token MergingZhijian Shu, Cheng Lin, Tao Xie, Wei Yin et al.CVPR 2026 · 17 citations
- FastVGGT: Fast Visual Geometry TransformerYou Shen, Zhipeng Zhang, Yansong Qu, Xiawu Zheng et al.ICLR 2026 · 73 citations
- Co-Me: Confidence Guided Token Merging for Visual Geometric TransformersYutian Chen, Yuheng Qiu, Ruogu Li, Jay Patrikar et al.CVPR 2026 · 9 citations
- QVGGT: Post-Training Quantized Visual Geometry Grounded TransformerZhizhen Pan, Hesong Wang, Huan WangCVPR 2026 · 2 citations
- FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor AttentionZipeng Wang, Dan XuCVPR 2026 · 14 citations
