CroMo: Cross-Modal Learning for Monocular Depth Estimation
Yannick Verdié, Jifei Song, Barnabé Mas, Benjamin Busam, Ales Leonardis, Steven McDonagh
Abstract
Learning-based depth estimation has witnessed recent progress in multiple directions; from self-supervision using monocular video to supervised methods offering highest accuracy. Complementary to supervision, further boosts to performance and robustness are gained by combining information from multiple signals. In this paper we systematically investigate key trade-offs associated with sensor and modality design choices as well as related model training strategies. Our study leads us to a new method, capable of connecting modality-specific advantages from polarisation, Time-of-Flight and structured-light inputs. We propose a novel pipeline capable of estimating depth from monocular polarisation for which we evaluate various training signals. The inversion of differentiable analytic models thereby connects scene geometry with polarisation and ToF signals and enables self-supervised and cross-modal learning. In the absence of existing multimodal datasets, we examine our approach with a custom-made multi-modal camera rig and collect CroMo; the first dataset to consist of synchronized stereo polarisation, indirect ToF and structured-light depth, captured at video rates. Extensive experiments on challenging video scenes confirm both qualitative and quantitative pipeline advantages where we are able to outperform competitive monocular depth estimation methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 77423b70-9f5d-468a-a3ee-50a62680b800Cited by top-tier papers7
- Rapid Network Adaptation: Learning to Adapt Neural Networks Using Test-Time FeedbackTeresa Yeo, Oguzhan Fatih Kar, Zahra Sodagar, Amir ZamirICCV 2023 · 10 citations
- DMR: Decomposed Multi-Modality Representations for Frames and Events Fusion in Visual Reinforcement LearningHaoran Xu, Peixi Peng, Guang Tan, Yuan Li et al.CVPR 2024 · 5 citations
- UnReflectAnything: RGB-Only Highlight Removal by Rendering Synthetic Specular SupervisionAlberto Rota, Mert Kiray, Mert Asim Karaoglu, Patrick Ruhkamp et al.CVPR 2026 · 2 citations
- Learning Accurate 3D Shape Based on Stereo Polarimetric ImagingTianyu Huang, Haoang Li, Kejing He, Congying Sui et al.CVPR 2023
- HouseCat6D - A Large-Scale Multi-Modal Category Level 6D Object Perception Dataset with Household Objects in Realistic ScenariosHyunJun Jung, Shun-Cheng Wu, Patrick Ruhkamp, Guangyao Zhai et al.CVPR 2024
Builds on8
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Free-Form Image Inpainting With Gated ConvolutionJiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen et al.ICCV 2019 · 1,990 citations
- Self-Supervised Monocular Depth HintsJamie Watson, Michael Firman, Gabriel J. Brostow, Daniyar TurmukhambetovICCV 2019 · 287 citations
- Learning an Augmented RGB Representation with Cross-Modal Knowledge Distillation for Action DetectionRui Dai, Srijan Das, François BrémondICCV 2021 · 50 citations
- Polarized Reflection Removal With Perfect Alignment in the WildChenyang Lei, Xuhua Huang, Mengdi Zhang, Qiong Yan et al.CVPR 2020
Related papers
- PolarDepth: Monocular Transparent Object Depth from Polar-Physics PriorsWen Dong, Haiyang Mei, Yinglian Ji, Zijun Zhang et al.ICML 2026
- DepthInSpace: Exploitation and Fusion of Multiple Video Frames for Structured-Light Depth EstimationMohammad Mahdi Johari, Camilla Carta, François FleuretICCV 2021 · 12 citations
- CROMA: Remote Sensing Representations with Contrastive Radar-Optical Masked AutoencodersAnthony Fuller, Koreen Millard, James R. GreenNeurIPS 2023 · 245 citations
- Multi-Modal Neural Radiance Field for Monocular Dense SLAM with a Light-Weight ToF SensorXinyang Liu, Yijin Li, Yanbin Teng, Hujun Bao et al.ICCV 2023 · 41 citations
- Two-in-One Depth: Bridging the Gap Between Monocular and Binocular Self-supervised Depth EstimationZhengming Zhou, Qiulei DongICCV 2023 · 16 citations
