Sound Adversarial Audio-Visual Navigation
Yinfeng Yu, Wenbing Huang, Fuchun Sun, Changan Chen, Yikai Wang, Xiaohong Liu
摘要
Audio-visual navigation task requires an agent to find a sound source in a realistic, unmapped 3D environment by utilizing egocentric audio-visual observations. Existing audio-visual navigation works assume a clean environment that solely contains the target sound, which, however, would not be suitable in most real-world applications due to the unexpected sound noise or intentional interference. In this work, we design an acoustically complex environment in which, besides the target sound, there exists a sound attacker playing a zero-sum game with the agent. More specifically, the attacker can move and change the volume and category of the sound to make the agent suffer from finding the sounding object while the agent tries to dodge the attack and navigate to the goal under the intervention. Under certain constraints to the attacker, we can improve the robustness of the agent towards unexpected sound attacks in audio-visual navigation. For better convergence, we develop a joint training mechanism by employing the property of a centralized critic with decentralized actors. Experiments on two real-world 3D scan datasets, Replica, and Matterport3D, verify the effectiveness and the robustness of the agent trained under our designed environment when transferred to the clean environment or the one containing sound attackers with random policy. Project: https://yyf17.github.io/SAAVN.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Find What You Want: Learning Demand-conditioned Object Attribute Space for Demand-driven NavigationHongcheng Wang, Andy Guan Hong Chen, Xiaoqi Li, Mingdong Wu 等NeurIPS 2023 · 被引用 33 次
- Modality-Independent Teachers Meet Weakly-Supervised Audio-Visual Event ParserYung-Hsuan Lai, Yen-Chun Chen, Frank WangNeurIPS 2023 · 被引用 27 次
- Omnidirectional Information Gathering for Knowledge Transfer-based Audio-Visual NavigationJinyu Chen, Wenguan Wang, Si Liu, Hongsheng Li 等ICCV 2023 · 被引用 21 次
- AdVerb: Visually Guided Audio DereverberationSanjoy Chowdhury, Sreyan Ghosh, Subhrajyoti Dasgupta, Anton Ratnarajah 等ICCV 2023 · 被引用 21 次
- MO-DDN: A Coarse-to-Fine Attribute-based Exploration Agent for Multi-Object Demand-driven NavigationHongcheng Wang, Peiqi Liu, Wenzhe Cai, Mingdong Wu 等NeurIPS 2024 · 被引用 12 次
它引用的顶会 Paper12
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- Learning To Explore Using Active Neural SLAMDevendra Singh Chaplot, Dhiraj Gandhi, Saurabh Gupta, Abhinav Gupta 等ICLR 2020 · 被引用 603 次
- Robust Deep Reinforcement Learning against Adversarial Perturbations on State ObservationsHuan Zhang, Hongge Chen, Chaowei Xiao, Bo Li 等NeurIPS 2020 · 被引用 437 次
- Adversarial Policies: Attacking Deep Reinforcement LearningAdam Gleave, Michael Dennis, Cody Wild, Neel Kant 等ICLR 2020 · 被引用 415 次
- Robust Reinforcement Learning on State Observations with Learned Optimal AdversaryHuan Zhang, Hongge Chen, Duane S. Boning, Cho-Jui HsiehICLR 2021 · 被引用 212 次
相关 Paper
- Learning to Set Waypoints for Audio-Visual NavigationChangan Chen, Sagnik Majumder, Ziad Al-Halah, Ruohan Gao 等ICLR 2021 · 被引用 28 次
- Semantic Audio-Visual Navigation in Continuous EnvironmentsYichen Zeng, Hebaixu Wang, Meng Liu, Yu Zhou 等CVPR 2026 · 被引用 1 次
- Semantic Audio-Visual NavigationChangan Chen, Ziad Al-Halah, Kristen GraumanCVPR 2021
- RobustNav: Towards Benchmarking Robustness in Embodied NavigationPrithvijit Chattopadhyay, Judy Hoffman, Roozbeh Mottaghi, Aniruddha KembhaviICCV 2021 · 被引用 68 次
- Towards Audio-Visual Navigation in Noisy Environments: A Large-Scale Benchmark Dataset and an Architecture Considering Multiple Sound-SourcesZhanbo Shi, Lin Zhang, Linfei Li, Ying ShenAAAI 2025 · 被引用 8 次
