Reconstruct, Inpaint, Test-Time Finetune: Dynamic Novel-view Synthesis from Monocular Videos
Kaihua Chen, Tarasha Khurana, Deva Ramanan
摘要
We explore novel-view synthesis for dynamic scenes from monocular videos. Prior approaches rely on costly test-time optimization of 4D representations or do not preserve scene geometry when trained in a feed-forward manner. Our approach is based on three key insights: (1) covisible pixels (that are visible in both the input and target views) can be rendered by first reconstructing the dynamic 3D scene and rendering the reconstruction from the novel-views and (2) hidden pixels in novel views can be"inpainted"with feed-forward 2D video diffusion models. Notably, our video inpainting diffusion model (CogNVS) can be self-supervised from 2D videos, allowing us to train it on a large corpus of in-the-wild videos. This in turn allows for (3) CogNVS to be applied zero-shot to novel test videos via test-time finetuning. We empirically verify that CogNVS outperforms almost all prior art for novel-view synthesis of dynamic scenes from monocular videos.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Any4D: Unified Feed-Forward Metric 4D ReconstructionJay Karhade, Nikhil Varma Keetha, Yuchen Zhang, Tanisha Gupta 等CVPR 2026 · 被引用 35 次
- ReDirector: Creating Any-Length Video Retakes with Rotary Camera EncodingByeongjun Park, Byung-Hoon Kim, Hyungjin Chung, Jong ChulCVPR 2026 · 被引用 10 次
- FreeOrbit4D: Training-Free Arbitrary Camera Redirection for Monocular Videos via Foreground-Complete 4D ReconstructionWei Cao, Hao Zhang, Fengrui Tian, Yulun Wu 等SIGGRAPH 2026 · 被引用 4 次
它引用的顶会 Paper67
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 被引用 4,089 次
- Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman 等ICCV 2021 · 被引用 2,700 次
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 被引用 2,336 次
相关 Paper
- Vivid4D: Improving 4D Reconstruction from Monocular Video by Video InpaintingJiaxin Huang, Sheng Miao, Bangbang Yang, Yuewen Ma 等ICCV 2025 · 被引用 2 次
- Dynamic View Synthesis as an Inverse ProblemHidir Yesiltepe, Pinar YanardagNeurIPS 2025 · 被引用 12 次
- Novel View Synthesis from A Few Glimpses via Test-Time Natural Video CompletionYan Xu, Yixing Wang, Stella X. YuNeurIPS 2025 · 被引用 4 次
- 4DGT: Learning a 4D Gaussian Transformer Using Real-World Monocular VideosZhen Xu, Zhengqin Li, Zhao Dong, Xiaowei Zhou 等NeurIPS 2025 · 被引用 51 次
- Guess The Unseen: Dynamic 3D Scene Reconstruction from Partial 2D GlimpsesInhee Lee, Byungjun Kim, Hanbyul JooCVPR 2024
