Multi-Camera Collaborative Depth Prediction via Consistent Structure Estimation
Jialei Xu, Xianming Liu, Yuanchao Bai, Junjun Jiang, Kaixuan Wang, Xiaozhi Chen, Xiangyang Ji
Abstract
Depth map estimation from images is an important task in robotic systems. Existing methods can be categorized into two groups including multi-view stereo and monocular depth estimation. The former requires cameras to have large overlapping areas and sufficient baseline between cameras, while the latter that processes each image independently can hardly guarantee the structure consistency between cameras. In this paper, we propose a novel multi-camera collaborative depth prediction method that does not require large overlapping areas while maintaining structure consistency between cameras. Specifically, we formulate the depth estimation as a weighted combination of depth basis, in which the weights are updated iteratively by a refinement network driven by the proposed consistency loss. During the iterative update, the results of depth estimation are compared across cameras and the information of overlapping areas is propagated to the whole depth maps with the help of basis formulation. Experimental results on DDAD and NuScenes datasets demonstrate the superior performance of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e9a1cd15-13fc-4c83-ae05-2d138889b419Cited by top-tier papers4
- R3D3: Dense 3D Reconstruction of Dynamic Scenes from Multiple CamerasAron Schmied, Tobias Fischer, Martin Danelljan, Marc Pollefeys et al.ICCV 2023 · 52 citations
- Diffusion-Augmented Depth Prediction with Sparse AnnotationsJiaqi Li, Yiran Wang, Zihao Huang, Jinghong Zheng et al.ACM MM 2023 · 9 citations
- Multi-Frame Self-Supervised Depth Estimation with Multi-Scale Feature Fusion in Dynamic ScenesJiquan Zhong, Xiaolin Huang, Xiao YuACM MM 2023 · 6 citations
- Sparse2DGS: Geometry-Prioritized Gaussian Splatting for Surface Reconstruction from Sparse ViewsJiang Wu, Rui Li, Yu Zhu, Rong Guo et al.CVPR 2025
Builds on13
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Enforcing Geometric Constraints of Virtual Normal for Depth PredictionWei Yin, Yifan Liu, Chunhua Shen, Youliang YanICCV 2019 · 487 citations
- nuScenes: A Multimodal Dataset for Autonomous DrivingHolger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora et al.CVPR 2020
- Cost Volume Pyramid Based Depth Inference for Multi-View StereoJiayu Yang, Wei Mao, José M. Álvarez, Miaomiao LiuCVPR 2020
Related papers
- Normal Assisted Stereo Depth EstimationUday Kusupati, Shuo Cheng, Rui Chen, Hao SuCVPR 2020
- MonoMVSNet: Monocular Priors Guided Multi-View Stereo NetworkJianfei Jiang, Qiankun Liu, Haochen Yu, Hongyuan Liu et al.ICCV 2025 · 3 citations
- Multi-view Depth Estimation using Epipolar Spatio-Temporal NetworksXiaoxiao Long, Lingjie Liu, Wei Li, Christian Theobalt et al.CVPR 2021
- PTC-Depth: Pose-Refined Monocular Depth Estimation with Temporal ConsistencyLeezy Han, Seunggyu Kim, Dongseok Shim, Hyeonbeom LeeCVPR 2026
- 3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene UnderstandingXiaoye Wang, Chen Tang, Xiangyu Yue, Wei-Hong LiCVPR 2026 · 2 citations
