Seeing With Sound: Long-Range Acoustic Beamforming for Multimodal Scene Understanding
Praneeth Chakravarthula, Jim Aldon D'Souza, Ethan Tseng, Joe Bartusek, Felix Heide
Abstract
Mobile robots, including autonomous vehicles rely heavily on sensors that use electromagnetic radiation like lidars, radars and cameras for perception. While effective in most scenarios, these sensors can be unreliable in unfavorable environmental conditions, including low-light scenarios and adverse weather, and they can only detect obstacles within their direct line-of-sight. Audible sound from other road users propagates as acoustic waves that carry information even in challenging scenarios. However, their low spatial resolution and lack of directional information have made them an overlooked sensing modality. In this work, we introduce long-range acoustic beamforming of sound produced by road users in-the-wild as a complementary sensing modality to traditional electromagnetic radiation-based sensors. To validate our approach and encourage further work in the field, we also introduce the first-ever multimodal long-range acoustic beamforming dataset. We propose a neural aperture expansion method for beamforming and demonstrate its effectiveness for multimodal automotive object detection when coupled with RGB images in challenging automotive scenarios, where camera-only approaches fail or are unable to provide ultra-fast acoustic sensing sampling rates. Data and code can be found here 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Segment beyond View: Handling Partially Missing Modality for Audio-Visual Semantic SegmentationRenjie Wu, Hu Wang, Feras Dayoub, Hsiang-Ting ChenAAAI 2024 · 11 citations
- LayoutFormer: Hierarchical Text Detection Towards Scene Text UnderstandingMin Liang, Jia-Wei Ma, Xiaobin Zhu, Jingyan Qin et al.CVPR 2024
Builds on4
- Self-Supervised Moving Vehicle Tracking With Stereo SoundChuang Gan, Hang Zhao, Peihao Chen, David D. Cox et al.ICCV 2019 · 157 citations
- Gated2Depth: Real-Time Dense Lidar From Gated ImagesTobias Gruber, Frank D. Julca-Aguilar, Mario Bijelic, Felix HeideICCV 2019 · 70 citations
- nuScenes: A Multimodal Dataset for Autonomous DrivingHolger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora et al.CVPR 2020
- Scalability in Perception for Autonomous Driving: Waymo Open DatasetPei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard et al.CVPR 2020
Related papers
- Robust Multimodal Vehicle Detection in Foggy Weather Using Complementary Lidar and Radar SignalsKun Qian, Shilin Zhu, Xinyu Zhang, Li Erran LiCVPR 2021
- Temporal and Spatial Representation Learning for Multimodal Low-Beam 3D Object DetectionLin Wang, Shiliang Sun, Jing ZhaoAAAI 2026
- Radar Fields: Frequency-Space Neural Scene Representations for FMCW RadarDavid Borts, Erich Liang, Tim Broedermann, Andrea Ramazzina et al.SIGGRAPH 2024 · 20 citations
- UltraSE: single-channel speech enhancement using ultrasoundKe Sun, Xinyu ZhangMobiCom 2021 · 69 citations
- Multimodal Neural Acoustic Fields for Immersive Mixed RealityGuaneen Tong, Johnathan Chi-Ho Leung, Xi Peng, Haosheng Shi et al.IEEE VR 2025 · 4 citations
