FlashDepth: Real-Time Streaming Video Depth Estimation at 2K Resolution
Gene Chou, Wenqi Xian, Guandao Yang, Mohamed Abdelfattah, Bharath Hariharan, Noah Snavely, Ning Yu, Paul Debevec
摘要
A versatile video depth estimation model should (1) be accurate and consistent across frames, (2) produce high-resolution depth maps, and (3) support real-time streaming. We propose FlashDepth, a method that satisfies all three requirements, performing depth estimation on a 2044x1148 streaming video at 24 FPS. We show that, with careful modifications to pretrained single-image depth models, these capabilities are enabled with relatively little data and training. We evaluate our approach across multiple unseen datasets against state-of-the-art depth models, and find that ours outperforms them in terms of boundary sharpness and speed by a significant margin, while maintaining competitive accuracy. We hope our model will enable various applications that require high-resolution depth, such as video editing, and online decision-making, such as robotics. We release all code and model weights at https://github.com/Eyeline-Research/FlashDepth
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Vista4D: Video Reshooting with 4D Point CloudsKuan Heng Lin, Zhizheng Liu, Pablo Salamanca, Yash Kant 等CVPR 2026 · 被引用 17 次
- DAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry EstimationTuan Duc Ngo, Jiahui Huang, Seoung Wug Oh, Kevin Blackburn-Matzen 等CVPR 2026 · 被引用 3 次
- Stabilizing Streaming Video Geometry via Dynamic Feature NormalizationXiaoyang Lyu, Muxin Liu, Xiaoshan Wu, Ruicheng Wang 等CVPR 2026 · 被引用 3 次
- Hyden: A Hybrid Dual-Path Encoder for Monocular Geometry of High-resolution ImagesZaiwei Zhang, Marc Mapeke, Wei Ye, Rakesh Ranjan 等ICLR 2026
它引用的顶会 Paper43
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
相关 Paper
- Neural Video Depth StabilizerYiran Wang, Min Shi, Jiaqi Li, Zihao Huang 等ICCV 2023
- Depth Any Video with Scalable Synthetic DataHonghui Yang, Di Huang, Wei Yin, Chunhua Shen 等ICLR 2025
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen 等SIGGRAPH 2020 · 被引用 321 次
- Learning Temporally Consistent Video Depth from Video Diffusion PriorsJiahao Shao, Yuanbo Yang, Hongyu Zhou, Youmin Zhang 等CVPR 2025
- Temporally Consistent Online Depth Estimation Using Point-Based FusionNumair Khan, Eric Penner, Douglas Lanman, Lei XiaoCVPR 2023
