AVLEN: Audio-Visual-Language Embodied Navigation in 3D Environments
Sudipta Paul, Amit Roy-Chowdhury, Anoop Cherian
摘要
Recent years have seen embodied visual navigation advance in two distinct directions: (i) in equipping the AI agent to follow natural language instructions, and (ii) in making the navigable world multimodal, e.g., audio-visual navigation. However, the real world is not only multimodal, but also often complex, and thus in spite of these advances, agents still need to understand the uncertainty in their actions and seek instructions to navigate. To this end, we present AVLEN -- an interactive agent for Audio-Visual-Language Embodied Navigation. Similar to audio-visual navigation tasks, the goal of our embodied agent is to localize an audio event via navigating the 3D visual world; however, the agent may also seek help from a human (oracle), where the assistance is provided in free-form natural language. To realize these abilities, AVLEN uses a multimodal hierarchical reinforcement learning backbone that learns: (a) high-level policies to choose either audio-cues for navigation or to query the oracle, and (b) lower-level policies to select navigation actions based on its audio-visual and language inputs. The policies are trained via rewarding for the success on the navigation task while minimizing the number of queries to the oracle. To empirically evaluate AVLEN, we present experiments on the SoundSpaces framework for semantic audio-visual navigation tasks. Our results show that equipping the agent to ask for help leads to a clear improvement in performance, especially in challenging cases, e.g., when the sound is unheard during training or in the presence of distractor sounds.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- CAVEN: An Embodied Conversational Agent for Efficient Audio-Visual Navigation in Noisy EnvironmentsXiulong Liu, Sudipta Paul, Moitreya Chatterjee, Anoop CherianAAAI 2024 · 被引用 16 次
- REGNav: Room Expert Guided Image-Goal NavigationPengna Li, Kangyi Wu, Jingwen Fu, Sanping ZhouAAAI 2025 · 被引用 15 次
- STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D IntelligenceZihan Liu, Zhikang Niu, Qiuyang Xiao, Zhisheng Zheng 等ICLR 2026 · 被引用 12 次
- RILA: Reflective and Imaginative Language Agent for Zero-Shot Semantic Audio-Visual NavigationZeyuan Yang, Jiageng Lin, Peihao Chen, Anoop Cherian 等CVPR 2024 · 被引用 5 次
- Learn How to See: Collaborative Embodied Learning for Object Detection and Camera AdjustingLingdong Shen, Chunlei Huo, Nuo Xu, Chaowei Han 等AAAI 2024 · 被引用 4 次
它引用的顶会 Paper9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- Vision-Language Navigation with Random Environmental MixupChong Liu, Fengda Zhu, Xiaojun Chang, Xiaodan Liang 等ICCV 2021 · 被引用 113 次
- Just Ask: An Interactive Learning Framework for Vision and Language NavigationTa-Chung Chi, Minmin Shen, Mihail Eric, Seokhwan Kim 等AAAI 2020 · 被引用 88 次
- Sub-Instruction Aware Vision-and-Language NavigationYicong Hong, Cristian Rodriguez Opazo, Qi Wu, Stephen GouldEMNLP 2020 · 被引用 55 次
相关 Paper
- NaVLA: A Vision-Language-Audio-Action Model for Multimodal Instruction NavigationJugang Fan, Peihao Chen, Changhao Li, Qing Du 等AAAI 2026
- VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street ViewRaphael Schumann, Wanrong Zhu, Weixi Feng, Tsu-Jui Fu 等AAAI 2024 · 被引用 122 次
- AdaNav: Adaptive Reasoning with Uncertainty for Vision-Language NavigationXin Ding, Jianyu Wei, Yifan Yang, Shiqi Jiang 等ICML 2026 · 被引用 6 次
- Learning to Set Waypoints for Audio-Visual NavigationChangan Chen, Sagnik Majumder, Ziad Al-Halah, Ruohan Gao 等ICLR 2021 · 被引用 28 次
- What You See Is What You Reach: Towards Spatial Navigation with High-Level Human InstructionsLingfeng Zhang, Haoxiang Fu, Xiaoshuai Hao, Shuyi Zhang 等AAAI 2026 · 被引用 2 次
