MoSca: Dynamic Gaussian Fusion from Casual Videos via 4D Motion Scaffolds
Jiahui Lei, Yijia Weng, Adam W. Harley, Leonidas J. Guibas, Kostas Daniilidis
Abstract
We introduce 4D Motion Scaffolds (MoSca), a modern 4D reconstruction system designed to reconstruct and synthesize novel views of dynamic scenes from monocular videos captured casually in the wild. To address such a challenging and ill-posed inverse problem, we leverage prior knowledge from foundational vision models and lift the video data to a novel Motion Scaffold (MoSca) representation, which compactly and smoothly encodes the underlying motions/deformations. The scene geometry and appearance are then disentangled from the deformation field and are encoded by globally fusing the Gaussians anchored onto the MoSca and optimized via Gaussian Splatting. Additionally, camera focal length and poses can be solved using bundle adjustment without the need of any other pose estimation tools. Experiments demonstrate state-of-the-art performance on dynamic rendering benchmarks and its effectiveness on real videos.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 916bc5ff-406d-405a-a967-b51932274990Cited by top-tier papers48
- PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video GenerationChen Wang, Chuhao Chen, Yiming Huang, Zhiyang Dou et al.NeurIPS 2025 · 50 citations
- GFlow: Recovering 4D World from Monocular VideoShizun Wang, Xingyi Yang, Qiuhong Shen, Zhenxiang Jiang et al.AAAI 2025 · 47 citations
- MoVieS: Motion-Aware 4D Dynamic View Synthesis in One SecondChenguo Lin, Yuchen Lin, Panwang Pan, Yifan Yu et al.CVPR 2026 · 38 citations
- Any4D: Unified Feed-Forward Metric 4D ReconstructionJay Karhade, Nikhil Varma Keetha, Yuchen Zhang, Tanisha Gupta et al.CVPR 2026 · 35 citations
- Shape of Motion: 4D Reconstruction From a Single VideoQianqian Wang, Vickie Ye, Hang Gao, Weijia Zeng et al.ICCV 2025 · 29 citations
Builds on46
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- Nerfies: Deformable Neural Radiance FieldsKeunhong Park, Utkarsh Sinha, Jonathan T. Barron, Sofien Bouaziz et al.ICCV 2021 · 1,442 citations
Related papers
- Motion Decoupled 3D Gaussian Splatting for Dynamic Object RepresentationXiao Hu, Libo Long, Jochen LangAAAI 2025 · 2 citations
- MotionScale: Reconstructing Appearance, Geometry, and Motion of Dynamic Scenes with Scalable 4D Gaussian SplattingHaoran Zhou, Gim Hee LeeCVPR 2026 · 1 citation
- MOSAIC-GS: Monocular Scene Reconstruction via Advanced Initialization for Complex Dynamic EnvironmentsSvitlana Morkva, Vaishakh Patil, Alessio Tonioni, Michael Oechsle et al.CVPR 2026
- MSCD-GS: Motion-Separated Cooperative Deblurring Dynamic Reconstruction via Gaussian Splattingyongjian liao, Xu Zou, Wenjun Chen, Huixuan Li et al.CVPR 2026
- 4D3R: Motion-Aware Neural Reconstruction and Rendering of Dynamic Scenes from Monocular VideosMengqi Guo, Bo Xu, Yanyan Li, Gim Hee LeeNeurIPS 2025 · 2 citations
