Consistent depth of moving objects in video
Zhoutong Zhang, Forrester Cole, Richard Tucker, William T. Freeman, Tali Dekel
Abstract
We present a method to estimate depth of a dynamic scene, containing arbitrary moving objects, from an ordinary video captured with a moving camera. We seek a geometrically and temporally consistent solution to this under-constrained problem: the depth predictions of corresponding points across frames should induce plausible, smooth motion in 3D. We formulate this objective in a new test-time training framework where a depth-prediction CNN is trained in tandem with an auxiliary scene-flow prediction MLP over the entire input video. By recursively unrolling the scene-flow prediction MLP over varying time steps, we compute both short-range scene flow to impose local smooth motion priors directly in 3D, and long-range scene flow to impose multi-view consistency constraints with wide baselines. We demonstrate accurate and temporally coherent results on a variety of challenging videos containing diverse moving objects (pets, people, cars), as well as camera motion. Our depth maps give rise to a number of depth-and-motion aware video editing effects such as object and lighting insertion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers39
- SceneScape: Text-Driven Consistent Scene GenerationRafail Fridman, Amit Abecasis, Yoni Kasten, Tali DekelNeurIPS 2023 · 196 citations
- Source-free Depth for Object Pop-outZongwei Wu, Danda Pani Paudel, Deng-Ping Fan, Jingjing Wang et al.ICCV 2023 · 110 citations
- Uncertainty-aware State Space Transformer for Egocentric 3D Hand Trajectory ForecastingWentao Bao, Lele Chen, Libing Zeng, Zhong Li et al.ICCV 2023 · 34 citations
- 4D Gaussian Splatting in the Wild with Uncertainty-Aware RegularizationMijeong Kim, Jongwoo Lim, Bohyung HanNeurIPS 2024 · 33 citations
- Pseudo-Generalized Dynamic View Synthesis from a VideoXiaoming Zhao, Alex Colburn, Fangchang Ma, Miguel Ángel Bautista et al.ICLR 2024 · 31 citations
Builds on8
- Nerfies: Deformable Neural Radiance FieldsKeunhong Park, Utkarsh Sinha, Jonathan T. Barron, Sofien Bouaziz et al.ICCV 2021 · 1,442 citations
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen et al.SIGGRAPH 2020 · 321 citations
- Occupancy Flow: 4D Reconstruction by Learning Particle DynamicsMichael Niemeyer, Lars M. Mescheder, Michael Oechsle, Andreas GeigerICCV 2019 · 314 citations
- Self-Supervised Learning With Geometric Constraints in Monocular Video: Connecting Flow, Depth, and CameraYuhua Chen, Cordelia Schmid, Cristian SminchisescuICCV 2019 · 265 citations
- Novel View Synthesis of Dynamic Scenes With Globally Coherent Depths From a Monocular CameraJae Shin Yoon, Kihwan Kim, Orazio Gallo, Hyun Soo Park et al.CVPR 2020
Related papers
- Self-Supervised Monocular Scene Flow EstimationJunhwa Hur, Stefan RothCVPR 2020
- Neural Scene Flow Fields for Space-Time View Synthesis of Dynamic ScenesZhengqi Li, Simon Niklaus, Noah Snavely, Oliver WangCVPR 2021
- Temporally Consistent Online Depth Estimation Using Point-Based FusionNumair Khan, Eric Penner, Douglas Lanman, Lei XiaoCVPR 2023
- Dynamo-Depth: Fixing Unsupervised Depth Estimation for Dynamical ScenesYihong Sun, Bharath HariharanNeurIPS 2023 · 58 citations
- Depth-Aware Test-Time Training for Zero-Shot Video Object SegmentationWeihuang Liu, Xi Shen, Haolun Li, Xiuli Bi et al.CVPR 2024
