Streaming Visual Geometry Transformer
Dong Zhuo, Wenzhao Zheng, Jiahe Guo, Yuqi Wu, Jie Zhou, Jiwen Lu
Abstract
Perceiving and reconstructing 3D geometry from videos is a fundamental yet challenging computer vision task. To facilitate interactive and low-latency applications, we propose a streaming visual geometry transformer that shares a similar philosophy with autoregressive large language models. We explore a simple and efficient design and employ a causal transformer architecture to process the input sequence in an online manner. We use temporal causal attention and cache the historical keys and values as implicit memory to enable efficient streaming long-term 3D reconstruction. This design can handle low-latency 3D reconstruction by incrementally integrating historical information while maintaining high-quality spatial consistency. For efficient training, we propose to distill knowledge from the dense bidirectional visual geometry grounded transformer (VGGT) to our causal model. For inference, our model supports the migration of optimized efficient attention operators (e.g., FlashAttention) from large language models. Extensive experiments on various 3D geometry perception benchmarks demonstrate that our model enhances inference speed in online scenarios while maintaining competitive performance, thereby facilitating scalable and interactive 3D vision systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d75786d-502f-414c-9319-2ab5a2073e7aCited by top-tier papers33
- Scal3R: Scalable Test-Time Training for Large-Scale 3D ReconstructionTao Xie, Peishan Yang, Yudong Jin, Yingfeng Cai et al.CVPR 2026 · 26 citations
- Gen3R: 3D Scene Generation Meets Feed-Forward ReconstructionJiaxin Huang, Yuanbo Yang, Bangbang Yang, Lin Ma et al.CVPR 2026 · 24 citations
- DVGT: Driving Visual Geometry TransformerSicheng Zuo, Zixun Xie, Wenzhao Zheng, Shaoqing Xu et al.CVPR 2026 · 23 citations
- SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal OverheadChaojun Ni, Chen Cheng, Xiaofeng Wang, Zheng Zhu et al.CVPR 2026 · 23 citations
- OmniVGGT: Omni-Modality Driven Visual Geometry Grounded TransformerHaosong Peng, Hao Li, Yalun Dai, Yushi Lan et al.CVPR 2026 · 22 citations
Builds on28
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Multiview Neural Surface Reconstruction by Disentangling Geometry and AppearanceLior Yariv, Yoni Kasten, Dror Moran, Meirav Galun et al.NeurIPS 2020 · 1,010 citations
- Common Objects in 3D: Large-Scale Learning and Evaluation of Real-life 3D Category ReconstructionJeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone et al.ICCV 2021 · 686 citations
- Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene UnderstandingMike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar et al.ICCV 2021 · 633 citations
Related papers
- STream3R: Scalable Sequential 3D Reconstruction with Causal TransformerYushi Lan, Yihang Luo, Fangzhou Hong, Shangchen Zhou et al.ICLR 2026 · 84 citations
- FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor AttentionZipeng Wang, Dan XuCVPR 2026 · 14 citations
- STAC: Plug-and-Play Spatio-Temporal Aware Cache Compression for Streaming 3D ReconstructionRunze Wang, Yuxuan Song, Youcheng Cai, Ligang LiuCVPR 2026 · 6 citations
- LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token MergingZhijian Shu, Cheng Lin, Tao Xie, Wei Yin et al.CVPR 2026 · 17 citations
- LongStream: Long-Sequence Streaming Autoregressive Visual GeometryChong Cheng, Xianda Chen, Tao Xie, Wei Yin et al.CVPR 2026 · 16 citations
