FlashDepth: Real-Time Streaming Video Depth Estimation at 2K Resolution
Gene Chou, Wenqi Xian, Guandao Yang, Mohamed Abdelfattah, Bharath Hariharan, Noah Snavely, Ning Yu, Paul Debevec
Abstract
A versatile video depth estimation model should (1) be accurate and consistent across frames, (2) produce high-resolution depth maps, and (3) support real-time streaming. We propose FlashDepth, a method that satisfies all three requirements, performing depth estimation on a 2044x1148 streaming video at 24 FPS. We show that, with careful modifications to pretrained single-image depth models, these capabilities are enabled with relatively little data and training. We evaluate our approach across multiple unseen datasets against state-of-the-art depth models, and find that ours outperforms them in terms of boundary sharpness and speed by a significant margin, while maintaining competitive accuracy. We hope our model will enable various applications that require high-resolution depth, such as video editing, and online decision-making, such as robotics. We release all code and model weights at https://github.com/Eyeline-Research/FlashDepth
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 19dcaf5a-a413-40f7-a656-46d7573d91f2Cited by top-tier papers4
- Vista4D: Video Reshooting with 4D Point CloudsKuan Heng Lin, Zhizheng Liu, Pablo Salamanca, Yash Kant et al.CVPR 2026 · 17 citations
- DAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry EstimationTuan Duc Ngo, Jiahui Huang, Seoung Wug Oh, Kevin Blackburn-Matzen et al.CVPR 2026 · 3 citations
- Stabilizing Streaming Video Geometry via Dynamic Feature NormalizationXiaoyang Lyu, Muxin Liu, Xiaoshan Wu, Ruicheng Wang et al.CVPR 2026 · 3 citations
- Hyden: A Hybrid Dual-Path Encoder for Monocular Geometry of High-resolution ImagesZaiwei Zhang, Marc Mapeke, Wei Ye, Rakesh Ranjan et al.ICLR 2026
Builds on43
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
Related papers
- Neural Video Depth StabilizerYiran Wang, Min Shi, Jiaqi Li, Zihao Huang et al.ICCV 2023
- Depth Any Video with Scalable Synthetic DataHonghui Yang, Di Huang, Wei Yin, Chunhua Shen et al.ICLR 2025
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen et al.SIGGRAPH 2020 · 321 citations
- Learning Temporally Consistent Video Depth from Video Diffusion PriorsJiahao Shao, Yuanbo Yang, Hongyu Zhou, Youmin Zhang et al.CVPR 2025
- Temporally Consistent Online Depth Estimation Using Point-Based FusionNumair Khan, Eric Penner, Douglas Lanman, Lei XiaoCVPR 2023
