Learning to Set Waypoints for Audio-Visual Navigation
Changan Chen, Sagnik Majumder, Ziad Al-Halah, Ruohan Gao, Santhosh Kumar Ramakrishnan, Kristen Grauman
Abstract
In audio-visual navigation, an agent intelligently travels through a complex, unmapped 3D environment using both sights and sounds to find a sound source (e.g., a phone ringing in another room). Existing models learn to act at a fixed granularity of agent motion and rely on simple recurrent aggregations of the audio observations. We introduce a reinforcement learning approach to audio-visual navigation with two key novel elements: 1) waypoints that are dynamically set and learned end-to-end within the navigation policy, and 2) an acoustic memory that provides a structured, spatially grounded record of what the agent has heard as it moves. Both new ideas capitalize on the synergy of audio and visual data for revealing the geometry of an unmapped space. We demonstrate our approach on two challenging datasets of real-world 3D scenes, Replica and Matterport3D. Our model improves the state of the art by a substantial margin, and our experiments reveal that learning the links between sights, sounds, and space is essential for audio-visual navigation. Project: http://vision.cs.utexas.edu/ projects/audio_visual_waypoints .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 211bc4a4-3ade-417b-8572-526e68012b4dCited by top-tier papers41
- Waypoint Models for Instruction-guided Navigation in Continuous EnvironmentsJacob Krantz, Aaron Gokaslan, Dhruv Batra, Stefan Lee et al.ICCV 2021 · 153 citations
- PONI: Potential Functions for ObjectGoal Navigation with Interaction-free LearningSanthosh Kumar Ramakrishnan, Devendra Singh Chaplot, Ziad Al-Halah, Jitendra Malik et al.CVPR 2022 · 129 citations
- RobustNav: Towards Benchmarking Robustness in Embodied NavigationPrithvijit Chattopadhyay, Judy Hoffman, Roozbeh Mottaghi, Aniruddha KembhaviICCV 2021 · 68 citations
- Toward Practical Monocular Indoor Depth EstimationCho-Ying Wu, Jialiang Wang, Michael Hall, Ulrich Neumann et al.CVPR 2022 · 68 citations
- Bridging the Gap Between Learning in Discrete and Continuous Environments for Vision-and-Language NavigationYicong Hong, Zun Wang, Qi Wu, Stephen GouldCVPR 2022 · 66 citations
Builds on5
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- Learning To Explore Using Active Neural SLAMDevendra Singh Chaplot, Dhiraj Gandhi, Saurabh Gupta, Abhinav Gupta et al.ICLR 2020 · 603 citations
- Hierarchical Foresight: Self-Supervised Learning of Long-Horizon Tasks via Visual Subgoal GenerationSuraj Nair, Chelsea FinnICLR 2020 · 152 citations
- Neural Topological SLAM for Visual NavigationDevendra Singh Chaplot, Ruslan Salakhutdinov, Abhinav Gupta, Saurabh GuptaCVPR 2020
- Ego-Topo: Environment Affordances From Egocentric VideoTushar Nagarajan, Yanghao Li, Christoph Feichtenhofer, Kristen GraumanCVPR 2020
Related papers
- Semantic Audio-Visual NavigationChangan Chen, Ziad Al-Halah, Kristen GraumanCVPR 2021
- Semantic Audio-Visual Navigation in Continuous EnvironmentsYichen Zeng, Hebaixu Wang, Meng Liu, Yu Zhou et al.CVPR 2026 · 1 citation
- Sound Adversarial Audio-Visual NavigationYinfeng Yu, Wenbing Huang, Fuchun Sun, Changan Chen et al.ICLR 2022 · 49 citations
- Zero Experience Required: Plug & Play Modular Transfer Learning for Semantic Visual NavigationZiad Al-Halah, Santhosh K. Ramakrishnan, Kristen GraumanCVPR 2022 · 52 citations
- AVLEN: Audio-Visual-Language Embodied Navigation in 3D EnvironmentsSudipta Paul, Amit Roy-Chowdhury, Anoop CherianNeurIPS 2022 · 43 citations
