Buffer Anytime: Zero-Shot Video Depth and Normal from Image Priors
Zhengfei Kuang, Tianyuan Zhang, Kai Zhang, Hao Tan, Sai Bi, Yiwei Hu, Zexiang Xu, Milos Hasan, Gordon Wetzstein, Fujun Luan
摘要
Time Time Input Frames Input Frames Figure 1. Buffer Anytime improves temporal consistency in video geometry estimation without paired training data. Top: Comparison of depth estimation between Depth Anything V2 [55] and our method on a challenging dynamic scene with lighting variations. While the original model shows inconsistent depth predictions across frames, our approach maintains stable depth estimates. Bottom: Surface normal estimation comparison between Marigold-E2E-FT [20] and our method on an outdoor scene with complex geometry. Our method preserves consistent normal maps across frames while maintaining accurate geometric details. In both cases, our method achieves better temporal consistency without requiring video-geometry paired training data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- NormalCrafter: Learning Temporally Consistent Normals from Video Diffusion PriorsYanrui Bin, Wenbo Hu, Haoyuan Wang, Xinya Chen 等ICCV 2025 · 被引用 27 次
- FlashDepth: Real-Time Streaming Video Depth Estimation at 2K ResolutionGene Chou, Wenqi Xian, Guandao Yang, Mohamed Abdelfattah 等ICCV 2025 · 被引用 1 次
- GemDepth: Geometry-Embedded Features for 3D-Consistent Video DepthYuecheng Liu, Junda Cheng, Longliang Liu, Wenjing Liao 等ICML 2026 · 被引用 1 次
- ST360D: Spatial-to-Temporal Consistency for Training-free 360 Monocular Depth EstimationZidong Cao, Jinjing Zhu, Hao Ai, Lutao Jiang 等NeurIPS 2025 · 被引用 1 次
- SCOPE: Scale-Consistent One-Pass Estimation of 3D GeometryZheng Zhang, Lihe Yang, Tianyu Yang, Chaohui Yu 等SIGGRAPH 2026
它引用的顶会 Paper23
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang 等ICLR 2024 · 被引用 1,493 次
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu 等CVPR 2024 · 被引用 847 次
相关 Paper
- DepthCrafter: Generating Consistent Long Depth Sequences for Open-world VideosWenbo Hu, Xiangjun Gao, Xiaoyu Li, Sijie Zhao 等CVPR 2025
- Depth Any Video with Scalable Synthetic DataHonghui Yang, Di Huang, Wei Yin, Chunhua Shen 等ICLR 2025
- Learning Temporally Consistent Video Depth from Video Diffusion PriorsJiahao Shao, Yuanbo Yang, Hongyu Zhou, Youmin Zhang 等CVPR 2025
- Stabilizing Streaming Video Geometry via Dynamic Feature NormalizationXiaoyang Lyu, Muxin Liu, Xiaoshan Wu, Ruicheng Wang 等CVPR 2026 · 被引用 3 次
- Less is More: Consistent Video Depth Estimation with Masked Frames ModelingYiran Wang, Zhiyu Pan, Xingyi Li, Zhiguo Cao 等ACM MM 2022 · 被引用 23 次
