OSN: Infinite Representations of Dynamic 3D Scenes from Monocular Videos
Ziyang Song, Jinxi Li, Bo Yang
Abstract
It has long been challenging to recover the underlying dynamic 3D scene representations from a monocular RGB video. Existing works formulate this problem into finding a single most plausible solution by adding various constraints such as depth priors and strong geometry constraints, ignoring the fact that there could be infinitely many 3D scene representations corresponding to a single dynamic video. In this paper, we aim to learn all plausible 3D scene configurations that match the input video, instead of just inferring a specific one. To achieve this ambitious goal, we introduce a new framework, called OSN. The key to our approach is a simple yet innovative object scale network together with a joint optimization module to learn an accurate scale range for every dynamic 3D object. This allows us to sample as many faithful 3D scene configurations as possible. Extensive experiments show that our method surpasses all baselines and achieves superior accuracy in dynamic novel view synthesis on multiple synthetic and real-world datasets. Most notably, our method demonstrates a clear advantage in learning fine-grained 3D scene geometry. Our code and data are available at https://github.com/vLAR-group/OSN
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Depth-supervised NeRF: Fewer Views and Faster Training for FreeKangle Deng, Andrew Liu, Jun-Yan Zhu, Deva RamananCVPR 2022 · 756 citations
- Non-Rigid Neural Radiance Fields: Reconstruction and Novel View Synthesis of a Dynamic Scene From Monocular VideoEdgar Tretschk, Ayush Tewari, Vladislav Golyanik, Michael Zollhöfer et al.ICCV 2021 · 617 citations
Related papers
- DGNS: Deformable Gaussian Splatting and Dynamic Neural Surface for Monocular Dynamic 3D ReconstructionXuesong Li, Jinguang Tong, Jie Hong, Vivien Rolland et al.ACM MM 2025
- Neural Scene Flow Fields for Space-Time View Synthesis of Dynamic ScenesZhengqi Li, Simon Niklaus, Noah Snavely, Oliver WangCVPR 2021
- Pseudo-Generalized Dynamic View Synthesis from a VideoXiaoming Zhao, Alex Colburn, Fangchang Ma, Miguel Ángel Bautista et al.ICLR 2024 · 31 citations
- Shape of Motion: 4D Reconstruction From a Single VideoQianqian Wang, Vickie Ye, Hang Gao, Weijia Zeng et al.ICCV 2025 · 29 citations
- MoDGS: Dynamic Gaussian Splatting from Casually-captured Monocular Videos with Depth PriorsQingming Liu, Yuan Liu, Jiepeng Wang, Xianqiang Lyu et al.ICLR 2025
