Towards Audio-Visual Navigation in Noisy Environments: A Large-Scale Benchmark Dataset and an Architecture Considering Multiple Sound-Sources
Zhanbo Shi, Lin Zhang, Linfei Li, Ying Shen
摘要
Audio-visual navigation has received considerable attention in recent years. However, the majority of related investigations have focused on single sound-source scenarios. Studies in this field for multiple sound-source scenarios remain underexplored due to the limitations of two aspects. First, the existing audio-visual navigation dataset only has limited audio samples, making it difficult to simulate diverse multiple sound-source environments. Second, existing navigation frameworks are mainly designed for single sound-source scenarios, thus their performance is severely reduced in multiple sound-source scenarios. In this work, we make an attempt to fill in these two research gaps to some extent. First, we establish a large-scale BEnchmark Dataset for Audio-VIsual Navigation, namely BeDAViN. This dataset consists of 2,258 audio samples with a total duration of 10.8 hours, which is more than 33 times longer than the existing audio dataset employed in the audio-visual navigation task. Second, we propose a new Embodied Navigation framework for MUltiple Sound-Sources Scenarios called ENMuS 3 . There are mainly two essential components in ENMuS 3 , the sound event descriptor and the multi-scale scene memory transformer. The former component equips the agent with the ability to extract spatial and semantic features of the target sound-source among multiple sound-sources, while the latter provides the ability to track the target object effectively in noisy environments. Experimental results on our BeDAViN show that ENMuS 3 strongly outperforms its counterparts with an orderof-magnitude improvement in success rates across diverse scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper10
- ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal EmbeddingsArjun Majumdar, Gunjan Aggarwal, Bhavika Devnani, Judy Hoffman 等NeurIPS 2022 · 被引用 344 次
- Hierarchical Object-to-Zone Graph for Object NavigationSixian Zhang, Xinhang Song, Yubing Bai, Weijie Li 等ICCV 2021 · 被引用 98 次
- Is Mapping Necessary for Realistic PointGoal Navigation?Ruslan Partsey, Erik Wijmans, Naoki Yokoyama, Oles Dobosevych 等CVPR 2022 · 被引用 36 次
- Learning to Set Waypoints for Audio-Visual NavigationChangan Chen, Sagnik Majumder, Ziad Al-Halah, Ruohan Gao 等ICLR 2021 · 被引用 28 次
- Omnidirectional Information Gathering for Knowledge Transfer-based Audio-Visual NavigationJinyu Chen, Wenguan Wang, Si Liu, Hongsheng Li 等ICCV 2023 · 被引用 21 次
相关 Paper
- NaVLA: A Vision-Language-Audio-Action Model for Multimodal Instruction NavigationJugang Fan, Peihao Chen, Changhao Li, Qing Du 等AAAI 2026
- Semantic Audio-Visual NavigationChangan Chen, Ziad Al-Halah, Kristen GraumanCVPR 2021
- CAVEN: An Embodied Conversational Agent for Efficient Audio-Visual Navigation in Noisy EnvironmentsXiulong Liu, Sudipta Paul, Moitreya Chatterjee, Anoop CherianAAAI 2024 · 被引用 16 次
- Towards Versatile Embodied NavigationHanqing Wang, Wei Liang, Luc Van Gool, Wenguan WangNeurIPS 2022 · 被引用 48 次
- Audio-Visual Instance SegmentationRuohao Guo, Xianghua Ying, Yaru Chen, Dantong Niu 等CVPR 2025
