Seeing With Sound: Long-Range Acoustic Beamforming for Multimodal Scene Understanding
Praneeth Chakravarthula, Jim Aldon D'Souza, Ethan Tseng, Joe Bartusek, Felix Heide
摘要
Mobile robots, including autonomous vehicles rely heavily on sensors that use electromagnetic radiation like lidars, radars and cameras for perception. While effective in most scenarios, these sensors can be unreliable in unfavorable environmental conditions, including low-light scenarios and adverse weather, and they can only detect obstacles within their direct line-of-sight. Audible sound from other road users propagates as acoustic waves that carry information even in challenging scenarios. However, their low spatial resolution and lack of directional information have made them an overlooked sensing modality. In this work, we introduce long-range acoustic beamforming of sound produced by road users in-the-wild as a complementary sensing modality to traditional electromagnetic radiation-based sensors. To validate our approach and encourage further work in the field, we also introduce the first-ever multimodal long-range acoustic beamforming dataset. We propose a neural aperture expansion method for beamforming and demonstrate its effectiveness for multimodal automotive object detection when coupled with RGB images in challenging automotive scenarios, where camera-only approaches fail or are unable to provide ultra-fast acoustic sensing sampling rates. Data and code can be found here 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Segment beyond View: Handling Partially Missing Modality for Audio-Visual Semantic SegmentationRenjie Wu, Hu Wang, Feras Dayoub, Hsiang-Ting ChenAAAI 2024 · 被引用 11 次
- LayoutFormer: Hierarchical Text Detection Towards Scene Text UnderstandingMin Liang, Jia-Wei Ma, Xiaobin Zhu, Jingyan Qin 等CVPR 2024
它引用的顶会 Paper4
- Self-Supervised Moving Vehicle Tracking With Stereo SoundChuang Gan, Hang Zhao, Peihao Chen, David D. Cox 等ICCV 2019 · 被引用 157 次
- Gated2Depth: Real-Time Dense Lidar From Gated ImagesTobias Gruber, Frank D. Julca-Aguilar, Mario Bijelic, Felix HeideICCV 2019 · 被引用 70 次
- nuScenes: A Multimodal Dataset for Autonomous DrivingHolger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora 等CVPR 2020
- Scalability in Perception for Autonomous Driving: Waymo Open DatasetPei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard 等CVPR 2020
相关 Paper
- Robust Multimodal Vehicle Detection in Foggy Weather Using Complementary Lidar and Radar SignalsKun Qian, Shilin Zhu, Xinyu Zhang, Li Erran LiCVPR 2021
- Temporal and Spatial Representation Learning for Multimodal Low-Beam 3D Object DetectionLin Wang, Shiliang Sun, Jing ZhaoAAAI 2026
- Radar Fields: Frequency-Space Neural Scene Representations for FMCW RadarDavid Borts, Erich Liang, Tim Broedermann, Andrea Ramazzina 等SIGGRAPH 2024 · 被引用 20 次
- UltraSE: single-channel speech enhancement using ultrasoundKe Sun, Xinyu ZhangMobiCom 2021 · 被引用 69 次
- Multimodal Neural Acoustic Fields for Immersive Mixed RealityGuaneen Tong, Johnathan Chi-Ho Leung, Xi Peng, Haosheng Shi 等IEEE VR 2025 · 被引用 4 次
