Self-supervised surround-view depth estimation with volumetric feature fusion
Jung-Hee Kim, Junhwa Hur, Tien Phuoc Nguyen, Seong-Gyun Jeong
Abstract
We present a self-supervised depth estimation approach using a unified volumetric feature fusion for surround-view images. Given a set of surround-view images, our method constructs a volumetric feature map by extracting image feature maps from surround-view images and fuse the feature maps into a shared, unified 3D voxel space. The volumetric feature map then can be used for estimating a depth map at each surround view by projecting it into an image coordinate. A volumetric feature contains 3D information at its local voxel coordinate; thus our method can also synthesize a depth map at arbitrary rotated viewpoints by projecting the volumetric feature map into the target viewpoints. Furthermore, assuming static camera extrinsics in the multi-camera system, we propose to estimate a canonical camera motion from the volumetric feature map. Our method leverages 3D spatiotemporal context to learn metric-scale depth and the canonical camera motion in a self-supervised manner. Our method outperforms the prior arts on DDAD and nuScenes datasets, especially estimating more accurate metric-scale depth and consistent depth between neighboring views. * denotes equal contribution. † This work has been done at 42dot Inc. 36th Conference on Neural Information Processing Systems (NeurIPS 2022). Depth estimation Pose decoder Depth decoder Input images Surround-view feature fusion Depth and motion decoding Volumetric feature encoder Canonical motion estimation
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Towards Zero-Shot Scale-Aware Monocular Depth EstimationVitor Guizilini, Igor Vasiljevic, Dian Chen, Rares Ambrus et al.ICCV 2023 · 129 citations
- DrivingForward: Feed-forward 3D Gaussian Splatting for Driving Scene Reconstruction from Flexible Surround-view InputQijian Tian, Xin Tan, Yuan Xie, Lizhuang MaAAAI 2025 · 45 citations
- GaussianOcc: Fully Self-Supervised and Efficient 3D Occupancy Estimation with Gaussian SplattingWanshui Gan, Fang Liu, Hongbin Xu, Ningkai Mo et al.ICCV 2025 · 7 citations
- PFDepth: Heterogeneous Pinhole-Fisheye Joint Depth Estimation via Distortion-aware Gaussian-Splatted Volumetric FusionZhiwei Zhang, Ruikai Xu, Weijian Zhang, Zhizhong Zhang et al.ACM MM 2025 · 2 citations
Builds on14
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- RegNeRF: Regularizing Neural Radiance Fields for View Synthesis from Sparse InputsMichael Niemeyer, Jonathan T. Barron, Ben Mildenhall, Mehdi S. M. Sajjadi et al.CVPR 2022 · 513 citations
- Is Pseudo-Lidar needed for Monocular 3D Object detection?Dennis Park, Rares Ambrus, Vitor Guizilini, Jie Li et al.ICCV 2021 · 404 citations
- FIERY: Future Instance Prediction in Bird's-Eye View from Surround Monocular CamerasAnthony Hu, Zak Murez, Nikhil Mohan, Sofía Dudas et al.ICCV 2021 · 329 citations
Related papers
- DLFusion: Painting-Depth Augmenting-LiDAR for Multimodal Fusion 3D Object DetectionJunyin Wang, Chenghu Du, Hui Li, Shengwu XiongACM MM 2023 · 3 citations
- SurroundOcc: Multi-Camera 3D Occupancy Prediction for Autonomous DrivingYi Wei, Linqing Zhao, Wenzhao Zheng, Zheng Zhu et al.ICCV 2023 · 380 citations
- FusionOcc: Multi-Modal Fusion for 3D Occupancy PredictionShuo Zhang, Yupeng Zhai, Jilin Mei, Yu HuACM MM 2024 · 5 citations
- OmniMVS: End-to-End Learning for Omnidirectional Stereo MatchingChanghee Won, Jongbin Ryu, Jongwoo LimICCV 2019 · 61 citations
- UniPAD: A Universal Pre-Training Paradigm for Autonomous DrivingHonghui Yang, Sha Zhang, Di Huang, Xiaoyang Wu et al.CVPR 2024 · 31 citations
