Instant4D: 4D Gaussian Splatting in Minutes
Zhanpeng Luo, Haoxi Ran, Li Lu
Abstract
Dynamic view synthesis has seen significant advances, yet reconstructing scenes from uncalibrated, casual video remains challenging due to slow optimization and complex parameter estimation. In this work, we present Instant4D, a monocular reconstruction system that leverages native 4D representation to efficiently process casual video sequences within minutes, without calibrated cameras or depth sensors. Our method begins with geometric recovery through deep visual SLAM, followed by grid pruning to optimize scene representation. Our design significantly reduces redundancy while maintaining geometric integrity, cutting model size to under 10% of its original footprint. To handle temporal dynamics efficiently, we introduce a streamlined 4D Gaussian representation, achieving a 30x speed-up and reducing training time to within two minutes, while maintaining competitive performance across several benchmarks. Our method reconstruct a single video within 10 minutes on the Dycheck dataset or for a typical 200-frame video. We further apply our model to in-the-wild videos, showcasing its generalizability. Our project website is published at https://instant4d.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5a7ed29b-b790-4776-ac6a-5e74e8e868d2Cited by top-tier papers3
- pySpatial: Generating 3D Visual Programs for Zero-Shot Spatial ReasoningZhanpeng Luo, Ce Zhang, Silong Yong, Cunxi Dai et al.ICLR 2026 · 15 citations
- FreeOrbit4D: Training-Free Arbitrary Camera Redirection for Monocular Videos via Foreground-Complete 4D ReconstructionWei Cao, Hao Zhang, Fengrui Tian, Yulun Wu et al.SIGGRAPH 2026 · 4 citations
- Space-Time Forecasting of Dynamic Scenes with Motion-aware Gaussian GroupingJunmyeong Lee, Hoseung Choi, Minsu ChoCVPR 2026 · 3 citations
Builds on26
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- Nerfies: Deformable Neural Radiance FieldsKeunhong Park, Utkarsh Sinha, Jonathan T. Barron, Sofien Bouaziz et al.ICCV 2021 · 1,442 citations
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
- Real-time Photorealistic Dynamic Scene Representation and Rendering with 4D Gaussian SplattingZeyu Yang, Hongye Yang, Zijie Pan, Li ZhangICLR 2024 · 529 citations
Related papers
- 4D-Fly: Fast 4D Reconstruction from a Single Monocular VideoDiankun Wu, Fangfu Liu, Yi-Hsin Hung, Yue Qian et al.CVPR 2025
- 4DGT: Learning a 4D Gaussian Transformer Using Real-World Monocular VideosZhen Xu, Zhengqin Li, Zhao Dong, Xiaowei Zhou et al.NeurIPS 2025 · 51 citations
- MoVieS: Motion-Aware 4D Dynamic View Synthesis in One SecondChenguo Lin, Yuchen Lin, Panwang Pan, Yifan Yu et al.CVPR 2026 · 38 citations
- ViDAR: Video Diffusion-Aware 4D Reconstruction From Monocular InputsMichal Nazarczuk, Sibi Catley-Chandar, Thomas Tanay, Zhensong Zhang et al.NeurIPS 2025 · 5 citations
- Orientation-anchored Hyper-Gaussian for 4D Reconstruction from Casual VideosJunyi Wu, Jiachen Tao, Haoxuan Wang, Gaowen Liu et al.NeurIPS 2025 · 9 citations
