STAC: Plug-and-Play Spatio-Temporal Aware Cache Compression for Streaming 3D Reconstruction
Runze Wang, Yuxuan Song, Youcheng Cai, Ligang Liu
Abstract
Online 3D reconstruction from streaming inputs requires both long-term temporal consistency and efficient memory usage. While causal VGGT transformers address this challenge through key-value (KV) cache mechanism, the linear growth of the cache introduces a significant memory bottleneck. When memory constraints trigger early eviction, reconstruction quality and temporal consistency deteriorate markedly. In this work, we observe that attention patterns in causal transformers for 3D reconstruction exhibit intrinsic spatio-temporal sparsity. Leveraging this insight, we propose STAC, a Spatio-Temporally Aware Cache compression framework specifically designed for streaming 3D reconstruction using large causal transformers. STAC incorporates three key components: a Working Temporal Token Caching mechanism that preserves long-term informative tokens based on decayed cumulative attention scores; a Long-term Spatial Token Caching scheme that consolidates spatially redundant tokens into voxel-aligned representations for memory-efficient storage; and a Chunk-based Multi-frame Optimization strategy that jointly optimizes consecutive frames to enhance temporal coherence and leverage GPU parallelism. Extensive experiments demonstrate that STAC achieves state-of-the-art reconstruction quality while reducing memory consumption by 8.5 and accelerating inference by a factor of 3.5, enabling scalable and real-time 3D reconstruction in streaming settings. The code will be made publicly available upon acceptance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on20
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii et al.CVPR 2024 · 302 citations
- Neural RGB-D Surface ReconstructionDejan Azinovic, Ricardo Martin-Brualla, Dan B. Goldman, Matthias Nießner et al.CVPR 2022 · 272 citations
Related papers
- Streaming Visual Geometry TransformerDong Zhuo, Wenzhao Zheng, Jiahe Guo, Yuqi Wu et al.ICLR 2026 · 109 citations
- STream3R: Scalable Sequential 3D Reconstruction with Causal TransformerYushi Lan, Yihang Luo, Fangzhou Hong, Shangchen Zhou et al.ICLR 2026 · 84 citations
- IncVGGT: Incremental VGGT for Memory-Bounded Long-Range 3D ReconstructionKeyu Fang, Changchun Zhou, Yuzhe Fu, Hai Li et al.ICLR 2026
- Accelerating Streaming Video Large Language Models via Hierarchical Token CompressionYiyu Wang, Xuyang Liu, Xiyan Gui, Xinying Lin et al.CVPR 2026 · 40 citations
- LongStream: Long-Sequence Streaming Autoregressive Visual GeometryChong Cheng, Xianda Chen, Tao Xie, Wei Yin et al.CVPR 2026 · 16 citations
