STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer
Yushi Lan, Yihang Luo, Fangzhou Hong, Shangchen Zhou, Honghua Chen, Zhaoyang Lyu, Bo Dai, Shuai Yang, Chen Change Loy, Xingang Pan
Abstract
We present STream3R, a novel approach to 3D reconstruction that reformulates pointmap prediction as a decoder-only Transformer problem. Existing state-of-the-art methods for multi-view reconstruction either depend on expensive global optimization or rely on simplistic memory mechanisms that scale poorly with sequence length. In contrast, STream3R introduces an streaming framework that processes image sequences efficiently using causal attention, inspired by advances in modern language modeling. By learning geometric priors from large-scale 3D datasets, STream3R generalizes well to diverse and challenging scenarios, including dynamic scenes where traditional methods often fail. Extensive experiments show that our method consistently outperforms prior work across both static and dynamic scene benchmarks. Moreover, STream3R is inherently compatible with LLM-style training infrastructure, enabling efficient large-scale pretraining and fine-tuning for various downstream 3D tasks. Our results underscore the potential of causal Transformer models for online 3D perception, paving the way for real-time 3D understanding in streaming environments. More details can be found in our project page: https://nirvanalan.github.io/projects/stream3r.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a136fdea-c499-4931-a46b-62260772ce3eCited by top-tier papers16
- Scal3R: Scalable Test-Time Training for Large-Scale 3D ReconstructionTao Xie, Peishan Yang, Yudong Jin, Yingfeng Cai et al.CVPR 2026 · 26 citations
- OmniVGGT: Omni-Modality Driven Visual Geometry Grounded TransformerHaosong Peng, Hao Li, Yalun Dai, Yushi Lan et al.CVPR 2026 · 22 citations
- AMB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with BackendHengyi Wang, Lourdes AgapitoCVPR 2026 · 17 citations
- LongStream: Long-Sequence Streaming Autoregressive Visual GeometryChong Cheng, Xianda Chen, Tao Xie, Wei Yin et al.CVPR 2026 · 16 citations
- 4RC: 4D Reconstruction via Conditional Querying Anytime and AnywhereYihang Luo, Shangchen Zhou, Yushi Lan, Xingang Pan et al.ICML 2026 · 12 citations
Builds on54
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
Related papers
- Streaming Visual Geometry TransformerDong Zhuo, Wenzhao Zheng, Jiahe Guo, Yuqi Wu et al.ICLR 2026 · 109 citations
- RnG: A Unified Transformer for Complete 3D Modeling from Partial ObservationsMochu Xiang, Zhelun Shen, Xuesong li, Jiahui Ren et al.CVPR 2026 · 2 citations
- STAC: Plug-and-Play Spatio-Temporal Aware Cache Compression for Streaming 3D ReconstructionRunze Wang, Yuxuan Song, Youcheng Cai, Ligang LiuCVPR 2026 · 6 citations
- LONG3R: Long Sequence Streaming 3D ReconstructionZhuoguang Chen, Minghui Qin, Tianyuan Yuan, Zhe Liu et al.ICCV 2025 · 3 citations
- GeoLRM: Geometry-Aware Large Reconstruction Model for High-Quality 3D Gaussian GenerationChubin Zhang, Hongliang Song, Yi Wei, Chen Yu et al.NeurIPS 2024 · 40 citations
