Continuous 3D Perception Model with Persistent State
Qianqian Wang, Yifei Zhang, Aleksander Holynski, Alexei A. Efros, Angjoo Kanazawa
2025Year
215Top-tier citations
Abstract
Static Scene Dynamic Scene Sparse Photo Collection Figure 1 . Continuous 3D Perception. Given a stream of RGB images as input, our approach enables dense 3D reconstruction in an online, continuous manner, estimating both camera parameters and dense 3D geometry with each incoming frame. Our framework supports various 3D tasks, processes inputs from video sequences and sparse photo collections, and can handle both static and dynamic scenes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4ab463a3-b3cd-496d-be94-d9f5382123d3Cited by top-tier papers215
- Depth Anything 3: Recovering the Visual Space from Any ViewsHaotong Lin, Sili Chen, Jun Hao Liew, Donny Y. Chen et al.ICLR 2026 · 720 citations
- π3: Permutation-Equivariant Visual Geometry LearningYifan Wang, Jianjun Zhou, Haoyi Zhu, Wenzheng Chang et al.ICLR 2026 · 318 citations
- VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) ManifoldDominic Maggio, Hyungtae Lim, Luca CarloneNeurIPS 2025 · 176 citations
- VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D ReconstructionZhiwen Fan, Jian Zhang, Renjie Li, Junge Zhang et al.CVPR 2026 · 171 citations
- Video World Models with Long-term Spatial MemoryTong Wu, Shuai Yang, Ryan Po, Yinghao Xu et al.NeurIPS 2025 · 145 citations
Builds on48
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
Related papers
- Memory-based Adapters for Online 3D Scene PerceptionXiuwei Xu, Chong Xia, Ziwei Wang, Linqing Zhao et al.CVPR 2024 · 4 citations
- ProDyG: Progressive Dynamic Scene Reconstruction via Gaussian Splatting from Monocular VideosShi Chen, Erik Sandström, Sandro Lombardi, Siyuan Li et al.NeurIPS 2025 · 1 citation
- Point4Cast: Streaming Dynamic Scene Reconstruction and ForecastingXinhang Liu, Pedro Miraldo, Suhas Lohit, Huaizu Jiang et al.CVPR 2026
- 4D Primitive-Mâché: Glueing Primitives for Persistent 4D Scene ReconstructionKirill Mazur, Marwan Taher, Andrew J. DavisonCVPR 2026 · 1 citation
- LivePose: Online 3D Reconstruction from Monocular Video with Dynamic Camera PosesNoah Stier, Baptiste Angles, Liang Yang, Yajie Yan et al.ICCV 2023 · 5 citations
