Jasmine: Harnessing Diffusion Prior for Self-supervised Depth Estimation
Jiyuan Wang, Chunyu Lin, Cheng Guan, Lang Nie, Jing He, Haodong Li, Kang Liao, Yao Zhao
摘要
In this paper, we propose Jasmine, the first Stable Diffusion (SD)-based self-supervised framework for monocular depth estimation, which effectively harnesses SD's visual priors to enhance the sharpness and generalization of unsupervised prediction. Previous SD-based methods are all supervised since adapting diffusion models for dense prediction requires high-precision supervision. In contrast, self-supervised reprojection suffers from inherent challenges (e.g., occlusions, texture-less regions, illumination variance), and the predictions exhibit blurs and artifacts that severely compromise SD's latent priors. To resolve this, we construct a novel surrogate task of hybrid image reconstruction. Without any additional supervision, it preserves the detail priors of SD models by reconstructing the images themselves while preventing depth estimation from degradation. Furthermore, to address the inherent misalignment between SD's scale and shift invariant estimation and self-supervised scale-invariant depth estimation, we build the Scale-Shift GRU. It not only bridges this distribution gap but also isolates the fine-grained texture of SD output against the interference of reprojection loss. Extensive experiments demonstrate that Jasmine achieves SoTA performance on the KITTI benchmark and exhibits superior zero-shot generalization across multiple datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Pixel-Perfect Depth with Semantics-Prompted Diffusion TransformersGangwei Xu, Haotong Lin, Hongcheng Luo, Xianqi Wang 等NeurIPS 2025 · 被引用 58 次
- Depth Any Panoramas: A Foundation Model for Panoramic Depth EstimationXin Lin, Meixi Song, Dizhe Zhang, Wenxuan Lu 等CVPR 2026 · 被引用 27 次
- DA2: Depth Anything in Any DirectionHaodong Li, Wangguandong Zheng, Jing He, Yuhao Liu 等ICLR 2026 · 被引用 23 次
- LongStream: Long-Sequence Streaming Autoregressive Visual GeometryChong Cheng, Xianda Chen, Tao Xie, Wei Yin 等CVPR 2026 · 被引用 16 次
- Semantic Context Matters: Improving Conditioning for Autoregressive ModelsDongyang Jin, Ryan Xu, Jianhao Zeng, Rui Lan 等CVPR 2026 · 被引用 12 次
它引用的顶会 Paper46
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 被引用 2,416 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
相关 Paper
- Repurposing Diffusion-Based Image Generators for Monocular Depth EstimationBingxin Ke, Anton Obukhov, Shengyu Huang, Nando Metzger 等CVPR 2024
- Exploiting Pseudo Labels in a Self-Supervised Learning Framework for Improved Monocular Depth EstimationAndra Petrovai, Sergiu NedevschiCVPR 2022 · 被引用 56 次
- Boost 3D Reconstruction Using Diffusion-Based Monocular Camera CalibrationJunyuan Deng, Wei Yin, Xiaoyang Guo, Qian Zhang 等ICCV 2025 · 被引用 2 次
- Iris: Bringing Real-World Priors into Diffusion Model for Monocular Depth EstimationXinhao Cai, Gensheng Pei, Zeren Sun, Yazhou Yao 等CVPR 2026 · 被引用 2 次
- Lotus: Diffusion-based Visual Foundation Model for High-quality Dense PredictionJing He, Haodong Li, Wei Yin, Yixun Liang 等ICLR 2025
