UniSplat: Unified Spatio-Temporal Fusion via 3D Latent Scaffolds for Dynamic Driving Scene Reconstruction
Chen Shi, Shaoshuai Shi, Xiaoyang Lyu, Chunyang Liu, Kehua Sheng, Bo Zhang, Li Jiang
Abstract
Feed-forward 3D reconstruction for autonomous driving has advanced rapidly, yet existing methods struggle with the joint challenges of sparse, non-overlapping camera views and complex scene dynamics. We present UniSplat, a general feed-forward framework that learns robust dynamic scene reconstruction through unified latent spatio-temporal fusion. UniSplat constructs a 3D latent scaffold, a structured representation that captures geometric and semantic scene context by leveraging pretrained foundation models. To effectively integrate information across spatial views and temporal frames, we introduce an efficient fusion mechanism that operates directly within the 3D scaffold, enabling consistent spatio-temporal alignment. To ensure complete and detailed reconstructions, we design a dual-branch decoder that generates dynamic-aware Gaussians from the fused scaffold by combining point-anchored refinement with voxel-based generation, and maintain a persistent memory of static Gaussians to enable streaming scene completion beyond current camera coverage. Extensive experiments on real-world datasets demonstrate that UniSplat achieves state-of-the-art performance in novel view synthesis, while providing robust and high-quality renderings even for viewpoints outside the original camera coverage.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2ce5293-8674-4455-a3a1-c878fc2b02ddCited by top-tier papers1
Ask how each one uses itBuilds on35
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- Point-NeRF: Point-based Neural Radiance FieldsQiangeng Xu, Zexiang Xu, Julien Philip, Sai Bi et al.CVPR 2022 · 510 citations
- Mip-Splatting: Alias-Free 3D Gaussian SplattingZehao Yu, Anpei Chen, Binbin Huang, Torsten Sattler et al.CVPR 2024 · 360 citations
- MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp DetailsRuicheng Wang, Sicheng Xu, Yue Dong, Yu Deng et al.NeurIPS 2025 · 308 citations
Related papers
- WorldSplat: Gaussian-Centric Feed-Forward 4D Scene Generation for Autonomous DrivingZiyue Zhu, Zhanqian Wu, Zhenxin Zhu, Lijun Zhou et al.ICLR 2026 · 11 citations
- DynamicVGGT: Learning Dynamic Point Maps for 4D Scene Reconstruction in Autonomous DrivingZhuolin He, Jing Li, Guanghao Li, Xiaolei Chen et al.CVPR 2026 · 5 citations
- StreamSplat: Towards Online Dynamic 3D Reconstruction from Uncalibrated Video StreamsZike Wu, Qi Yan, Xuanyu Yi, Lele Wang et al.ICLR 2026 · 9 citations
- Diff4Splat: Repurposing Video Diffusion Models for Dynamic Scene GenerationPanwang Pan, Chenguo Lin, Chenxin Li, Jingjing Zhao et al.CVPR 2026
- Learning 3D Representations for Spatial Intelligence from Unposed Multi-View ImagesBo Zhou, Qiuxia Lai, Zeren Sun, Xiangbo Shu et al.CVPR 2026 · 1 citation
