Test-Time Prompt Tuning for Zero-Shot Depth Completion
Chanhwi Jeong, Inhwan Bae, Jin-Hwi Park, Hae-Gon Jeon
摘要
Least squares fitting (c) Our fine-tuning (d) Our visual prompting (a) Monocular depth estimation
Input RGB Initial Depth Metric-scale sparse depth
Learnable Parameters: 0.9M Learnable Parameters: 363M Figure 1. Comparison of methods for correcting the inaccurate scale of initial depth estimates produced by the depth foundation model using sensor-derived sparse depth with an accurate scale, represented as yellow dots. (a) The depth foundation model generates an initial depth map from a single RGB input. (b) The least squares fitting aligns the initial depth map with the sparse depth [54]. (c) The parameters of the foundation model are fine-tuned under supervision from the sparse depth. (d) Alternatively, the model's parameters are frozen, and a visual prompt is learned to adjust the depth estimates. Please refer to Fig. 3 for depth profiling results and error analysis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Large Depth Completion Model from Sparse ObservationsZhu Yu, zhengyi zhao, Runmin Zhang, Lingteng Qiu 等ICLR 2026 · 被引用 8 次
- When CLIP Sees More, It Fights Back Harder: Multi-View Guided Adaptive Counterattacks for Test-Time Adversarial RobustnessSunoh Kim, Daeho UmCVPR 2026 · 被引用 2 次
- ORCaS: Unsupervised Depth Completion via Occluded Region Completion as SupervisionHyoungseob Park, Runjian Chen, Patrick Rim, Dong Lao 等ICLR 2026
- Entropy-Monitored Kernelized Token Distillation for Audio-Visual CompressionHyoungseob Park, Lipeng Ke, Pritish Mohapatra, Huajun Ying 等ICLR 2026
- SS-TPT: Stability and Suitability-Guided Test-Time Prompt Tuning for Adversarially Robust Vision-Language ModelsSunoh Kim, Daeho UmICML 2026
它引用的顶会 Paper35
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
相关 Paper
- Prompting Depth Anything for 4K Resolution Accurate Metric Depth EstimationHaotong Lin, Sida Peng, Jingxiao Chen, Songyou Peng 等CVPR 2025
- Depth Prompting for Sensor-Agnostic Depth EstimationJin-Hwi Park, Chanhwi Jeong, Junoh Lee, Hae-Gon JeonCVPR 2024 · 被引用 4 次
- RPG360: Robust 360 Depth Estimation with Perspective Foundation Models and Graph OptimizationDongki Jung, Jaehoon Choi, Yonghan Lee, Dinesh ManochaNeurIPS 2025 · 被引用 4 次
- Radar-Guided Polynomial Fitting for Metric Depth EstimationPatrick Rim, Hyoungseob Park, Vadim Ezhov, Jeffrey Moon 等CVPR 2026 · 被引用 7 次
- Diving into the Fusion of Monocular Priors for Generalized Stereo MatchingChengtang Yao, Lidong Yu, Zhidan Liu, Jiaxi Zeng 等ICCV 2025 · 被引用 3 次
