Pseudo-Generalized Dynamic View Synthesis from a Video
Xiaoming Zhao, Alex Colburn, Fangchang Ma, Miguel Ángel Bautista, Joshua M. Susskind, Alexander G. Schwing
Abstract
Rendering scenes observed in a monocular video from novel viewpoints is a challenging problem. For static scenes the community has studied both scene-specific optimization techniques, which optimize on every test scene, and generalized techniques, which only run a deep net forward pass on a test scene. In contrast, for dynamic scenes, scene-specific optimization techniques exist, but, to our best knowledge, there is currently no generalized method for dynamic novel view synthesis from a given monocular video. To explore whether generalized dynamic novel view synthesis from monocular videos is possible today, we establish an analysis framework based on existing techniques and work toward the generalized approach. We find a pseudo-generalized process without scene-specific appearance optimization is possible, but geometrically and temporally consistent depth estimates are needed. Despite no scene-specific appearance optimization, the pseudogeneralized approach improves upon some scene-specific methods. For more information see project page at https://xiaoming-zhao.github.io/projects/pgdvs . * Work done as part of an internship at Apple. 1 "Generalized" is defined as no need of optimization / fitting / training / fine-tuning on test scenes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers20
- Feed-Forward Bullet-Time Reconstruction of Dynamic Scenes from Monocular VideosHanxue Liang, Jiawei Ren, Ashkan Mirzaei, Antonio Torralba et al.NeurIPS 2025 · 52 citations
- MoVieS: Motion-Aware 4D Dynamic View Synthesis in One SecondChenguo Lin, Yuchen Lin, Panwang Pan, Yifan Yu et al.CVPR 2026 · 38 citations
- GoMAvatar: Efficient Animatable Human Modeling from Monocular Video Using Gaussians-on-MeshJing Wen, Xiaoming Zhao, Zhongzheng Ren, Alexander G. Schwing et al.CVPR 2024 · 33 citations
- 4D-LRM: Large Space-Time Reconstruction Model From and To Any View at Any TimeZiqiao Ma, Xuweiyi Chen, Shoubin Yu, Sai Bi et al.NeurIPS 2025 · 15 citations
- DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular VideosChieh Hubert Lin, Zhaoyang Lv, Songyin Wu, Zhen Xu et al.NeurIPS 2025 · 15 citations
Builds on47
- Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman et al.ICCV 2021 · 2,700 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan et al.CVPR 2022 · 1,603 citations
- MVSNeRF: Fast Generalizable Radiance Field Reconstruction from Multi-View StereoAnpei Chen, Zexiang Xu, Fuqiang Zhao, Xiaoshuai Zhang et al.ICCV 2021 · 1,024 citations
- Zip-NeRF: Anti-Aliased Grid-Based Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan et al.ICCV 2023 · 799 citations
Related papers
- MoDGS: Dynamic Gaussian Splatting from Casually-captured Monocular Videos with Depth PriorsQingming Liu, Yuan Liu, Jiepeng Wang, Xianqiang Lyu et al.ICLR 2025
- Reconstruct, Inpaint, Test-Time Finetune: Dynamic Novel-view Synthesis from Monocular VideosKaihua Chen, Tarasha Khurana, Deva RamananNeurIPS 2025 · 17 citations
- OSN: Infinite Representations of Dynamic 3D Scenes from Monocular VideosZiyang Song, Jinxi Li, Bo YangICML 2024
- Neural Scene Flow Fields for Space-Time View Synthesis of Dynamic ScenesZhengqi Li, Simon Niklaus, Noah Snavely, Oliver WangCVPR 2021
- Guess The Unseen: Dynamic 3D Scene Reconstruction from Partial 2D GlimpsesInhee Lee, Byungjun Kim, Hanbyul JooCVPR 2024
