Semantic Audio-Visual Navigation
Changan Chen, Ziad Al-Halah, Kristen Grauman
Abstract
Recent work on audio-visual navigation assumes a constantly-sounding target and restricts the role of audio to signaling the target's position. We introduce semantic audio-visual navigation, where objects in the environment make sounds consistent with their semantic meaning (e.g., toilet flushing, door creaking) and acoustic events are sporadic or short in duration. We propose a transformer-based model to tackle this new semantic AudioGoal task, incorporating an inferred goal descriptor that captures both spatial and semantic properties of the target. Our model's persistent multimodal memory enables it to reach the goal even long after the acoustic event stops. In support of the new task, we also expand the SoundSpaces audio simulations to provide semantically grounded sounds for an array of objects in Matterport3D. Our method strongly outperforms existing audio-visual navigation methods by learning to associate semantic, acoustic, and visual cues. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75ad09c7-0d6d-4173-84bb-144d1828b96cCited by top-tier papers39
- 🏘️ ProcTHOR: Large-Scale Embodied AI Using Procedural GenerationMatt Deitke, Eli VanderBilt, Alvaro Herrasti, Luca Weihs et al.NeurIPS 2022 · 596 citations
- PONI: Potential Functions for ObjectGoal Navigation with Interaction-free LearningSanthosh Kumar Ramakrishnan, Devendra Singh Chaplot, Ziad Al-Halah, Jitendra Malik et al.CVPR 2022 · 129 citations
- Few-Shot Audio-Visual Learning of Environment AcousticsSagnik Majumder, Changan Chen, Ziad Al-Halah, Kristen GraumanNeurIPS 2022 · 80 citations
- AV-NeRF: Learning Neural Fields for Real-World Audio-Visual Scene SynthesisSusan Liang, Chao Huang, Yapeng Tian, Anurag Kumar et al.NeurIPS 2023 · 77 citations
- Toward Practical Monocular Indoor Depth EstimationCho-Ying Wu, Jialiang Wang, Michael Hall, Ulrich Neumann et al.CVPR 2022 · 68 citations
Builds on9
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- Object Goal Navigation using Goal-Oriented Semantic ExplorationDevendra Singh Chaplot, Dhiraj Gandhi, Abhinav Gupta, Ruslan SalakhutdinovNeurIPS 2020 · 857 citations
- DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion FramesErik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee et al.ICLR 2020 · 608 citations
- Learning To Explore Using Active Neural SLAMDevendra Singh Chaplot, Dhiraj Gandhi, Saurabh Gupta, Abhinav Gupta et al.ICLR 2020 · 603 citations
- Co-Separating Sounds of Visual ObjectsRuohan Gao, Kristen GraumanICCV 2019 · 224 citations
Related papers
- Semantic Audio-Visual Navigation in Continuous EnvironmentsYichen Zeng, Hebaixu Wang, Meng Liu, Yu Zhou et al.CVPR 2026 · 1 citation
- Learning to Set Waypoints for Audio-Visual NavigationChangan Chen, Sagnik Majumder, Ziad Al-Halah, Ruohan Gao et al.ICLR 2021 · 28 citations
- Visual Acoustic MatchingChangan Chen, Ruohan Gao, Paul Calamia, Kristen GraumanCVPR 2022 · 42 citations
- Towards Audio-Visual Navigation in Noisy Environments: A Large-Scale Benchmark Dataset and an Architecture Considering Multiple Sound-SourcesZhanbo Shi, Lin Zhang, Linfei Li, Ying ShenAAAI 2025 · 8 citations
- NaVLA: A Vision-Language-Audio-Action Model for Multimodal Instruction NavigationJugang Fan, Peihao Chen, Changhao Li, Qing Du et al.AAAI 2026
