Towards Audio-Visual Navigation in Noisy Environments: A Large-Scale Benchmark Dataset and an Architecture Considering Multiple Sound-Sources
Zhanbo Shi, Lin Zhang, Linfei Li, Ying Shen
Abstract
Audio-visual navigation has received considerable attention in recent years. However, the majority of related investigations have focused on single sound-source scenarios. Studies in this field for multiple sound-source scenarios remain underexplored due to the limitations of two aspects. First, the existing audio-visual navigation dataset only has limited audio samples, making it difficult to simulate diverse multiple sound-source environments. Second, existing navigation frameworks are mainly designed for single sound-source scenarios, thus their performance is severely reduced in multiple sound-source scenarios. In this work, we make an attempt to fill in these two research gaps to some extent. First, we establish a large-scale BEnchmark Dataset for Audio-VIsual Navigation, namely BeDAViN. This dataset consists of 2,258 audio samples with a total duration of 10.8 hours, which is more than 33 times longer than the existing audio dataset employed in the audio-visual navigation task. Second, we propose a new Embodied Navigation framework for MUltiple Sound-Sources Scenarios called ENMuS 3 . There are mainly two essential components in ENMuS 3 , the sound event descriptor and the multi-scale scene memory transformer. The former component equips the agent with the ability to extract spatial and semantic features of the target sound-source among multiple sound-sources, while the latter provides the ability to track the target object effectively in noisy environments. Experimental results on our BeDAViN show that ENMuS 3 strongly outperforms its counterparts with an orderof-magnitude improvement in success rates across diverse scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 90995cbf-cb8b-40dc-8e16-ca305d670ac5Cited by top-tier papers1
Ask how each one uses itBuilds on10
- ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal EmbeddingsArjun Majumdar, Gunjan Aggarwal, Bhavika Devnani, Judy Hoffman et al.NeurIPS 2022 · 344 citations
- Hierarchical Object-to-Zone Graph for Object NavigationSixian Zhang, Xinhang Song, Yubing Bai, Weijie Li et al.ICCV 2021 · 98 citations
- Is Mapping Necessary for Realistic PointGoal Navigation?Ruslan Partsey, Erik Wijmans, Naoki Yokoyama, Oles Dobosevych et al.CVPR 2022 · 36 citations
- Learning to Set Waypoints for Audio-Visual NavigationChangan Chen, Sagnik Majumder, Ziad Al-Halah, Ruohan Gao et al.ICLR 2021 · 28 citations
- Omnidirectional Information Gathering for Knowledge Transfer-based Audio-Visual NavigationJinyu Chen, Wenguan Wang, Si Liu, Hongsheng Li et al.ICCV 2023 · 21 citations
Related papers
- NaVLA: A Vision-Language-Audio-Action Model for Multimodal Instruction NavigationJugang Fan, Peihao Chen, Changhao Li, Qing Du et al.AAAI 2026
- Semantic Audio-Visual NavigationChangan Chen, Ziad Al-Halah, Kristen GraumanCVPR 2021
- CAVEN: An Embodied Conversational Agent for Efficient Audio-Visual Navigation in Noisy EnvironmentsXiulong Liu, Sudipta Paul, Moitreya Chatterjee, Anoop CherianAAAI 2024 · 16 citations
- Towards Versatile Embodied NavigationHanqing Wang, Wei Liang, Luc Van Gool, Wenguan WangNeurIPS 2022 · 48 citations
- Audio-Visual Instance SegmentationRuohao Guo, Xianghua Ying, Yaru Chen, Dantong Niu et al.CVPR 2025
