Buffer Anytime: Zero-Shot Video Depth and Normal from Image Priors
Zhengfei Kuang, Tianyuan Zhang, Kai Zhang, Hao Tan, Sai Bi, Yiwei Hu, Zexiang Xu, Milos Hasan, Gordon Wetzstein, Fujun Luan
Abstract
Time Time Input Frames Input Frames Figure 1. Buffer Anytime improves temporal consistency in video geometry estimation without paired training data. Top: Comparison of depth estimation between Depth Anything V2 [55] and our method on a challenging dynamic scene with lighting variations. While the original model shows inconsistent depth predictions across frames, our approach maintains stable depth estimates. Bottom: Surface normal estimation comparison between Marigold-E2E-FT [20] and our method on an outdoor scene with complex geometry. Our method preserves consistent normal maps across frames while maintaining accurate geometric details. In both cases, our method achieves better temporal consistency without requiring video-geometry paired training data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- NormalCrafter: Learning Temporally Consistent Normals from Video Diffusion PriorsYanrui Bin, Wenbo Hu, Haoyuan Wang, Xinya Chen et al.ICCV 2025 · 27 citations
- FlashDepth: Real-Time Streaming Video Depth Estimation at 2K ResolutionGene Chou, Wenqi Xian, Guandao Yang, Mohamed Abdelfattah et al.ICCV 2025 · 1 citation
- GemDepth: Geometry-Embedded Features for 3D-Consistent Video DepthYuecheng Liu, Junda Cheng, Longliang Liu, Wenjing Liao et al.ICML 2026 · 1 citation
- ST360D: Spatial-to-Temporal Consistency for Training-free 360 Monocular Depth EstimationZidong Cao, Jinjing Zhu, Hao Ai, Lutao Jiang et al.NeurIPS 2025 · 1 citation
- SCOPE: Scale-Consistent One-Pass Estimation of 3D GeometryZheng Zhang, Lihe Yang, Tianyu Yang, Chaohui Yu et al.SIGGRAPH 2026
Builds on23
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang et al.ICLR 2024 · 1,493 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
Related papers
- DepthCrafter: Generating Consistent Long Depth Sequences for Open-world VideosWenbo Hu, Xiangjun Gao, Xiaoyu Li, Sijie Zhao et al.CVPR 2025
- Depth Any Video with Scalable Synthetic DataHonghui Yang, Di Huang, Wei Yin, Chunhua Shen et al.ICLR 2025
- Learning Temporally Consistent Video Depth from Video Diffusion PriorsJiahao Shao, Yuanbo Yang, Hongyu Zhou, Youmin Zhang et al.CVPR 2025
- Stabilizing Streaming Video Geometry via Dynamic Feature NormalizationXiaoyang Lyu, Muxin Liu, Xiaoshan Wu, Ruicheng Wang et al.CVPR 2026 · 3 citations
- Less is More: Consistent Video Depth Estimation with Masked Frames ModelingYiran Wang, Zhiyu Pan, Xingyi Li, Zhiguo Cao et al.ACM MM 2022 · 23 citations
