DCDepth: Progressive Monocular Depth Estimation in Discrete Cosine Domain
Kun Wang, Zhiqiang Yan, Junkai Fan, Wanlu Zhu, Xiang Li, Jun Li, Jian Yang
摘要
In this paper, we introduce DCDepth, a novel framework for the long-standing monocular depth estimation task. Moving beyond conventional pixel-wise depth estimation in the spatial domain, our approach estimates the frequency coefficients of depth patches after transforming them into the discrete cosine domain. This unique formulation allows for the modeling of local depth correlations within each patch. Crucially, the frequency transformation segregates the depth information into various frequency components, with low-frequency components encapsulating the core scene structure and high-frequency components detailing the finer aspects. This decomposition forms the basis of our progressive strategy, which begins with the prediction of low-frequency components to establish a global scene context, followed by successive refinement of local details through the prediction of higher-frequency components. We conduct comprehensive experiments on NYU-Depth-V2, TOFDC, and KITTI datasets, and demonstrate the state-of-the-art performance of DCDepth. Code is available at https://github.com/w2kun/DCDepth.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Hyper-Depth: Hypergraph-Based Multi-Scale Representation Fusion for Monocular Depth EstimationLin Bie, Siqi Li, Yifan Feng, Yue GaoICCV 2025
- Scalable Autoregressive Monocular Depth EstimationJinhong Wang, Jian Liu, Dongqi Tang, Weiqiang Wang 等CVPR 2025
- MER-Tracker: Towards High-Speed 3D Point Tracking via Multi-View Event-RGB Hybrid CamerasYiqian Chang, Qinghong Ye, Haoran Xu, Jianing Li 等CVPR 2026
- Object Concepts Emerge from MotionHaoqian Liang, Xiaohui Wang, Zhichao Li, Ya Yang 等NeurIPS 2025
- DORNet: A Degradation Oriented and Regularized Network for Blind Depth Super-ResolutionZhengxue Wang, Zhiqiang Yan, Jinshan Pan, Guangwei Gao 等CVPR 2025
它引用的顶会 Paper23
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Enforcing Geometric Constraints of Virtual Normal for Depth PredictionWei Yin, Yifan Liu, Chunhua Shen, Youliang YanICCV 2019 · 被引用 487 次
- Neural Window Fully-connected CRFs for Monocular Depth EstimationWeihao Yuan, Xiaodong Gu, Zuozhuo Dai, Siyu Zhu 等CVPR 2022 · 被引用 320 次
相关 Paper
- Patch-Wise Attention Network for Monocular Depth EstimationSihaeng Lee, Janghyeon Lee, Byungju Kim, Eojindl Yi 等AAAI 2021 · 被引用 84 次
- Transformer-Based Attention Networks for Continuous Pixel-Wise PredictionGuanglei Yang, Hao Tang, Mingli Ding, Nicu Sebe 等ICCV 2021 · 被引用 246 次
- R-MSFM: Recurrent Multi-Scale Feature Modulation for Monocular Depth EstimatingZhongkai Zhou, Xinnan Fan, Pengfei Shi, Yuanxue XinICCV 2021 · 被引用 150 次
- Single Image Depth Prediction With Wavelet DecompositionMichaël Ramamonjisoa, Michael Firman, Jamie Watson, Vincent Lepetit 等CVPR 2021
- MonoDETR: Depth-guided Transformer for Monocular 3D Object DetectionRenrui Zhang, Han Qiu, Tai Wang, Ziyu Guo 等ICCV 2023 · 被引用 175 次
