Beyond Image to Depth: Improving Depth Prediction Using Echoes
Kranti Kumar Parida, Siddharth Srivastava, Gaurav Sharma
Abstract
We address the problem of estimating depth with multi modal audio visual data. Inspired by the ability of animals, such as bats and dolphins, to infer distance of objects with echolocation, some recent methods have utilized echoes for depth estimation. We propose an end-to-end deep learning based pipeline utilizing RGB images, binaural echoes and estimated material properties of various objects within a scene. We argue that the relation between image, echoes and depth, for different scene elements, is greatly influenced by the properties of those elements, and a method designed to leverage this information can lead to significantly improved depth estimation from audio visual inputs. We propose a novel multi modal fusion technique, which incorporates the material properties explicitly while combining audio (echoes) and visual modalities to predict the scene depth. We show empirically, with experiments on Replica dataset, that the proposed method obtains 28% improvement in RMSE compared to the state-of-the-art audio-visual depth prediction method. To demonstrate the effectiveness of our method on larger dataset, we report competitive performance on Matterport3D, proposing to use it as a multimodal depth prediction benchmark with echoes for the first time. We also analyse the proposed method with exhaustive ablation experiments and qualitative results. The code and models are available at https://krantiparida . github.io/projects/bimgdepth.html
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a92f0843-1fb6-4dbc-9e96-00491d993dd9Cited by top-tier papers6
- Sound Localization from Motion: Jointly Learning Sound Direction and Camera RotationZiyang Chen, Shengyi Qian, Andrew OwensICCV 2023 · 21 citations
- Dense 2D-3D Indoor Prediction with Sound via Aligned Cross-Modal DistillationHeeseung Yun, Joonil Na, Gunhee KimICCV 2023 · 8 citations
- EchoDiffusion: Waveform Conditioned Diffusion Models for Echo-Based Depth EstimationWenjie Zhang, Jun Yin, Long Ma, Peng Yu et al.AAAI 2025 · 2 citations
- Few-shot Acoustic Synthesis with Multimodal Flow MatchingAmandine BrunettoCVPR 2026 · 2 citations
- Supervising Sound Localization by In-the-wild EgomotionAnna Min, Ziyang Chen, Hang Zhao, Andrew OwensCVPR 2025
Builds on8
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- Pseudo-LiDAR++: Accurate Depth for 3D Object Detection in Autonomous DrivingYurong You, Yan Wang, Wei-Lun Chao, Divyansh Garg et al.ICLR 2020 · 439 citations
- The Sound of MotionsHang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio TorralbaICCV 2019 · 271 citations
- Co-Separating Sounds of Visual ObjectsRuohan Gao, Kristen GraumanICCV 2019 · 224 citations
Related papers
- Towards Multimodal Depth Estimation from Light FieldsTitus Leistner, Radek Mackowiak, Lynton Ardizzone, Ullrich Köthe et al.CVPR 2022 · 14 citations
- GLAVNet: Global-Local Audio-Visual Cues for Fine-Grained Material RecognitionFengmin Shi, Jie Guo, Haonan Zhang, Shan Yang et al.CVPR 2021
- Scene-Aware Audio Rendering via Deep Acoustic AnalysisZhenyu Tang, Nicholas J. Bryan, Dingzeyu Li, Timothy R. Langlois et al.IEEE VR 2020 · 37 citations
- There Is More Than Meets the Eye: Self-Supervised Multi-Object Detection and Tracking With Sound by Distilling Multimodal KnowledgeFrancisco Rivera Valverde, Juana Valeria Hurtado, Abhinav ValadaCVPR 2021
- Binaural Audio-Visual LocalizationXinyi Wu, Zhenyao Wu, Lili Ju, Song WangAAAI 2021 · 32 citations
