Vivid4D: Improving 4D Reconstruction from Monocular Video by Video Inpainting
Jiaxin Huang, Sheng Miao, Bangbang Yang, Yuewen Ma, Yiyi Liao
Abstract
Reconstructing 4D dynamic scenes from casually captured monocular videos is valuable but highly challenging, as each timestamp is observed from a single viewpoint. We introduce Vivid4D, a novel approach that enhances 4D monocular video synthesis by augmenting observation views - synthesizing multi-view videos from a monocular input. Unlike existing methods that either solely leverage geometric priors for supervision or use generative priors while overlooking geometry, we integrate both. This reformulates view augmentation as a video inpainting task, where observed views are warped into new viewpoints based on monocular depth priors. To achieve this, we train a video inpainting model on unposed web videos with synthetically generated masks that mimic warping occlusions, ensuring spatially and temporally consistent completion of missing regions. To further mitigate inaccuracies in monocular depth priors, we introduce an iterative view augmentation strategy and a robust reconstruction loss. Experiments demonstrate that our method effectively improves monocular 4D scene reconstruction and completion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2ccee36e-edf5-49a9-a438-518dad4d34ccCited by top-tier papers4
- NeoVerse: Enhancing 4D World Model with in-the-wild Monocular VideosYuxue Yang, Lue Fan, Ziqi Shi, Junran Peng et al.CVPR 2026 · 42 citations
- EditCtrl: Disentangled Local and Global Control for Real-Time Generative Video EditingYehonathan Litman, Shikun Liu, Dario Seyb, Nicholas Milef et al.CVPR 2026 · 5 citations
- GP-4DGS: Probabilistic 4D Gaussian Splatting from Monocular Video via Variational Gaussian ProcessesMijeong Kim, Jungtaek Kim, Bohyung HanCVPR 2026 · 3 citations
- Pantheon360: Taming Digital Twin Generation via 3D-Aware 360° Video DiffusionTing-Hsuan Chen, Ying-Huan Chen, Tao Tu, Jie-Ying Lee et al.CVPR 2026 · 2 citations
Builds on61
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman et al.ICCV 2021 · 2,700 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
Related papers
- Scaling4D: Pushing the Frontier of Video Novel View Synthesis through Large-Scale Monocular VideosHongrui Cai, Junjie Luo, Zhihong Fu, Shengnan Zhu et al.CVPR 2026
- EasyCreator: Empowering 4D Creation through Video InpaintingYue Ma, Kunyu Feng, Xinhua Zhang, Hongyu Liu et al.ICLR 2026 · 47 citations
- Motion 3-to-4: 3D Motion Reconstruction for 4D SynthesisHongyuan Chen, Xingyu Chen, Zexiang Xu, Anpei ChenCVPR 2026 · 17 citations
- Video Camera Trajectory Editing with Generative Rendering from Estimated GeometryJunyoung Seo, Jisang Han, Jaewoo Jung, Siyoon Jin et al.AAAI 2026
- MoVieS: Motion-Aware 4D Dynamic View Synthesis in One SecondChenguo Lin, Yuchen Lin, Panwang Pan, Yifan Yu et al.CVPR 2026 · 38 citations
