Deep 3D Mask Volume for View Synthesis of Dynamic Scenes
Kai-En Lin, Lei Xiao, Feng Liu, Guowei Yang, Ravi Ramamoorthi
Abstract
Image view synthesis has seen great success in reconstructing photorealistic visuals, thanks to deep learning and various novel representations. The next key step in immersive virtual experiences is view synthesis of dynamic scenes. However, several challenges exist due to the lack of high-quality training datasets, and the additional time dimension for videos of dynamic scenes. To address this issue, we introduce a multi-view video dataset, captured with a custom 10-camera rig in 120FPS. The dataset contains 96 high-quality scenes showing various visual effects and human interactions in outdoor scenes. We develop a new algorithm, Deep 3D Mask Volume, which enables temporally-stable view extrapolation from binocular videos of dynamic scenes, captured by static cameras. Our algorithm addresses the temporal inconsistency of disocclusions by identifying the error-prone areas with a 3D mask volume, and replaces them with static background observed throughout the video. Our method enables manipulation in 3D space as opposed to simple 2D masks, We demonstrate better temporal stability than frame-by-frame static view synthesis methods, or those that use 2D masks. The resulting view synthesis videos show minimal flickering artifacts and allow for larger translational movements.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 49ed8863-62fe-43ac-8a0b-30b657b66f6cCited by top-tier papers14
- DynPoint: Dynamic Neural Point For View SynthesisKaichen Zhou, Jia-Xing Zhong, Sangyun Shin, Kai Lu et al.NeurIPS 2023 · 46 citations
- 3D Moments from Near-Duplicate PhotosQianqian Wang, Zhengqi Li, David Salesin, Noah Snavely et al.CVPR 2022 · 17 citations
- Replay: Multi-modal Multi-view Acted Videos for Casual HolographyRoman Shapovalov, Yanir Kleiman, Ignacio Rocco, David Novotný et al.ICCV 2023 · 11 citations
- Factorized Motion Fields for Fast Sparse Input Dynamic View SynthesisNagabhushan Somraj, Kapil Choudhary, Sai Harsha Mupparaju, Rajiv SoundararajanSIGGRAPH 2024 · 6 citations
- DiVa-360: The Dynamic Visual Dataset for Immersive Neural FieldsCheng-You Lu, Peisen Zhou, Angela Xing, Chandradeep Pokhariya et al.CVPR 2024 · 4 citations
Builds on6
- Immersive light field video with a layered mesh representationMichael Broxton, John Flynn, Ryan S. Overbeck, Daniel Erickson et al.SIGGRAPH 2020 · 271 citations
- IBRNet: Learning Multi-View Image-Based RenderingQianqian Wang, Zhicheng Wang, Kyle Genova, Pratul P. Srinivasan et al.CVPR 2021
- Local Implicit Grid Representations for 3D ScenesChiyu Max Jiang, Avneesh Sud, Ameesh Makadia, Jingwei Huang et al.CVPR 2020
- Novel View Synthesis of Dynamic Scenes With Globally Coherent Depths From a Monocular CameraJae Shin Yoon, Kihwan Kim, Orazio Gallo, Hyun Soo Park et al.CVPR 2020
- 4D Visualization of Dynamic Events From Unconstrained Multi-View VideosAayush Bansal, Minh Vo, Yaser Sheikh, Deva Ramanan et al.CVPR 2020
Related papers
- DynamicStereo: Consistent Dynamic Depth from Stereo VideosNikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova et al.CVPR 2023
- Vivid4D: Improving 4D Reconstruction from Monocular Video by Video InpaintingJiaxin Huang, Sheng Miao, Bangbang Yang, Yuewen Ma et al.ICCV 2025 · 2 citations
- NVFi: Neural Velocity Fields for 3D Physics Learning from Dynamic VideosJinxi Li, Ziyang Song, Bo YangNeurIPS 2023 · 39 citations
- DreamScene4D: Dynamic Multi-Object Scene Generation from Monocular VideosWen-Hsuan Chu, Lei Ke, Katerina FragkiadakiNeurIPS 2024 · 75 citations
- MPI-Flow: Learning Realistic Optical Flow with Multiplane ImagesYingping Liang, Jiaming Liu, Debing Zhang, Ying FuICCV 2023 · 12 citations
