Monocular Dynamic View Synthesis: A Reality Check
Hang Gao, Ruilong Li, Shubham Tulsiani, Bryan Russell, Angjoo Kanazawa
Abstract
We study the recent progress on dynamic view synthesis (DVS) from monocular video. Though existing approaches have demonstrated impressive results, we show a discrepancy between the practical capture process and the existing experimental protocols, which effectively leaks in multi-view signals during training. We define effective multi-view factors (EMFs) to quantify the amount of multi-view signal present in the input capture sequence based on the relative camera-scene motion. We introduce two new metrics: co-visibility masked image metrics and correspondence accuracy, which overcome the issue in existing protocols. We also propose a new iPhone dataset that includes more diverse real-life deformation sequences. Using our proposed experimental protocol, we show that the state-of-the-art approaches observe a 1-2 dB drop in masked PSNR in the absence of multi-view cues and 4-5 dB drop when modeling complex motion. Code and data can be found at https://hangg7.com/dycheck .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers114
- NeRFPlayer: A Streamable Dynamic Scene Representation with Decomposed Neural Radiance FieldsLiangchen Song, Anpei Chen, Zhong Li, Zhang Chen et al.IEEE VR 2023 · 246 citations
- Tracking Everything Everywhere All at OnceQianqian Wang, Yen-Yu Chang, Ruojin Cai, Zhengqi Li et al.ICCV 2023 · 238 citations
- Consistent4D: Consistent 360° Dynamic Object Generation from Monocular VideoYanqin Jiang, Li Zhang, Jin Gao, Weiming Hu et al.ICLR 2024 · 120 citations
- NerfAcc: Efficient Sampling Accelerates NeRFsRuilong Li, Hang Gao, Matthew Tancik, Angjoo KanazawaICCV 2023 · 119 citations
- Nerfbusters: Removing Ghostly Artifacts from Casually Captured NeRFsFrederik Warburg, Ethan Weber, Matthew Tancik, Aleksander Holynski et al.ICCV 2023 · 92 citations
Builds on24
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan et al.CVPR 2022 · 1,603 citations
- Non-Rigid Neural Radiance Fields: Reconstruction and Novel View Synthesis of a Dynamic Scene From Monocular VideoEdgar Tretschk, Ayush Tewari, Vladislav Golyanik, Michael Zollhöfer et al.ICCV 2021 · 617 citations
- Dynamic View Synthesis from Dynamic Monocular VideoChen Gao, Ayush Saraf, Johannes Kopf, Jia-Bin HuangICCV 2021 · 522 citations
- HumanNeRF: Free-viewpoint Rendering of Moving People from Monocular VideoChung-Yi Weng, Brian Curless, Pratul P. Srinivasan, Jonathan T. Barron et al.CVPR 2022 · 411 citations
Related papers
- Deep 3D Mask Volume for View Synthesis of Dynamic ScenesKai-En Lin, Lei Xiao, Feng Liu, Guowei Yang et al.ICCV 2021 · 42 citations
- Scaling4D: Pushing the Frontier of Video Novel View Synthesis through Large-Scale Monocular VideosHongrui Cai, Junjie Luo, Zhihong Fu, Shengnan Zhu et al.CVPR 2026
- Enhanced Stable View SynthesisNishant Jain, Suryansh Kumar, Luc Van GoolCVPR 2023
- Novel View Synthesis with View-Dependent Effects from a Single ImageJuan Luis Gonzalez Bello, Munchurl KimCVPR 2024
- WildRayZer: Self-supervised Large View Synthesis in Dynamic EnvironmentsXuweiyi Chen, Wentao Zhou, Zezhou ChengCVPR 2026 · 5 citations
