Less is More: Consistent Video Depth Estimation with Masked Frames Modeling
Yiran Wang, Zhiyu Pan, Xingyi Li, Zhiguo Cao, Ke Xian, Jianming Zhang
Abstract
Temporal consistency is the key challenge of video depth estimation. Previous works are based on additional optical flow or camera poses, which is time-consuming. By contrast, we derive consistency with less information. Since videos inherently exist with heavy temporal redundancy, a missing frame could be recovered from neighboring ones. Inspired by this, we propose the frame masking network (FM-Net), a spatial-temporal transformer network predicting the depth of masked frames based on their neighboring frames. By reconstructing masked temporal features, the FMNet can learn intrinsic inter-frame correlations, which leads to consistency. Compared with prior arts, experimental results demonstrate that our approach achieves comparable spatial accuracy and higher temporal consistency without any additional information. Our work provides a new perspective on consistent video depth estimation. Our official project page is https://github.com/RaymondWang987/FMNet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b588604-108f-4fe3-868b-8fcbc0ca035cCited by top-tier papers20
- MAMo: Leveraging Memory and Attention for Monocular Video Depth EstimationRajeev Yasarla, Hong Cai, Jisoo Jeong, Yunxiao Shi et al.ICCV 2023 · 31 citations
- Constraining Depth Map Geometry for Multi-View Stereo: A Dual-Depth Approach with Saddle-shaped Depth CellsXinyi Ye, Weiyue Zhao, Tianqi Liu, Zihao Huang et al.ICCV 2023 · 29 citations
- Self-Guided Masked AutoencoderJeongwoo Shin, Inseo Lee, Junho Lee, Joonseok LeeNeurIPS 2024 · 18 citations
- Diffusion-Augmented Depth Prediction with Sparse AnnotationsJiaqi Li, Yiran Wang, Zihao Huang, Jinghong Zheng et al.ACM MM 2023 · 9 citations
- A robust inlier identification algorithm for point cloud registration via 𝓁0-minimizationYinuo Jiang, Xiuchuan Tang, Cheng Cheng, Ye YuanNeurIPS 2024 · 5 citations
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
Related papers
- GemDepth: Geometry-Embedded Features for 3D-Consistent Video DepthYuecheng Liu, Junda Cheng, Longliang Liu, Wenjing Liao et al.ICML 2026 · 1 citation
- Multi-view Depth Estimation using Epipolar Spatio-Temporal NetworksXiaoxiao Long, Lingjie Liu, Wei Li, Christian Theobalt et al.CVPR 2021
- Learning Structure Affinity for Video Depth EstimationYuanzhouhan Cao, Yidong Li, Haokui Zhang, Chao Ren et al.ACM MM 2021 · 12 citations
- DLVINet: Advancing Dual-Lens Video Inpainting Beyond Parallax ConstraintsZhiliang Wu, Kun Li, Yunqiu Xu, Hehe Fan et al.AAAI 2026 · 1 citation
- Video Semantic Segmentation via Sparse Temporal TransformerJiangtong Li, Wentao Wang, Junjie Chen, Li Niu et al.ACM MM 2021 · 47 citations
