Neural Video Depth Stabilizer
Yiran Wang, Min Shi, Jiaqi Li, Zihao Huang, Zhiguo Cao, Jianming Zhang, Ke Xian, Guosheng Lin
摘要
Video depth estimation aims to infer temporally consistent depth. Some methods achieve temporal consistency by finetuning a single-image depth model during test time using geometry and re-projection constraints, which is inefficient and not robust. An alternative approach is to learn how to enforce temporal consistency from data, but this requires well-designed models and sufficient video depth data. To address these challenges, we propose a plug-and-play framework called Neural Video Depth Stabilizer (NVDS) that stabilizes inconsistent depth estimations and can be applied to different single-image depth models without extra effort. We also introduce a large-scale dataset, Video Depth in the Wild (VDW), which consists of 14,203 videos with over two million frames, making it the largest natural-scene video depth dataset to our knowledge. We evaluate our method on the VDW dataset as well as two public benchmarks and demonstrate significant improvements in consistency, accuracy, and efficiency compared to previous approaches. Our work serves as a solid baseline and provides a data foundation for learning-based video depth models. We will release our dataset and code for future research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper42
- STream3R: Scalable Sequential 3D Reconstruction with Causal TransformerYushi Lan, Yihang Luo, Fangzhou Hong, Shangchen Zhou 等ICLR 2026 · 被引用 84 次
- OmniVDiff: Omni Controllable Video Diffusion for Generation and UnderstandingDianbing Xi, Jiepeng Wang, Yuanzhi Liang, Xi Qiu 等AAAI 2026 · 被引用 14 次
- 4RC: 4D Reconstruction via Conditional Querying Anytime and AnywhereYihang Luo, Shangchen Zhou, Yushi Lan, Xingang Pan 等ICML 2026 · 被引用 12 次
- Light-X: Generative 4D Video Rendering with Camera and Illumination ControlTianqi Liu, Zhaoxi Chen, Zihao Huang, Shaocong Xu 等ICLR 2026 · 被引用 12 次
- Diffusion-Augmented Depth Prediction with Sparse AnnotationsJiaqi Li, Yiran Wang, Zihao Huang, Jinghong Zheng 等ACM MM 2023 · 被引用 9 次
它引用的顶会 Paper19
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- GMFlow: Learning Optical Flow via Global MatchingHaofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi 等CVPR 2022 · 被引用 353 次
相关 Paper
- FlashDepth: Real-Time Streaming Video Depth Estimation at 2K ResolutionGene Chou, Wenqi Xian, Guandao Yang, Mohamed Abdelfattah 等ICCV 2025 · 被引用 1 次
- Depth Any Video with Scalable Synthetic DataHonghui Yang, Di Huang, Wei Yin, Chunhua Shen 等ICLR 2025
- 3D Video Stabilization With Depth Estimation by CNN-Based OptimizationYao-Chih Lee, Kuan-Wei Tseng, Yu-Ta Chen, Chien-Cheng Chen 等CVPR 2021
- DepthSync: Diffusion Guidance-Based Depth Synchronization for Scale- and Geometry-Consistent Video Depth EstimationYuejiang Dong, Wang Zhao, Jiale Xu, Ying Shan 等ICCV 2025
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen 等SIGGRAPH 2020 · 被引用 321 次
