Move2Hear: Active Audio-Visual Source Separation
Sagnik Majumder, Ziad Al-Halah, Kristen Grauman
Abstract
We introduce the active audio-visual source separation problem, where an agent must move intelligently in order to better isolate the sounds coming from an object of interest in its environment. The agent hears multiple audio sources simultaneously (e.g., a person speaking down the hall in a noisy household) and it must use its eyes and ears to automatically separate out the sounds originating from a target object within a limited time budget. Towards this goal, we introduce a reinforcement learning approach that trains movement policies controlling the agent's camera and microphone placement over time, guided by the improvement in predicted audio separation quality. We demonstrate our approach in scenarios motivated by both augmented reality (system is already co-located with the target object) and mobile robotics (agent begins arbitrarily far from the target object). Using state-of-the-art realistic audio-visual simulations in 3D environments, we demonstrate our model's ability to find minimal movement sequences with maximal payoff for audio source separation. Project: http://vision. cs.utexas.edu/projects/move2hear .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Few-Shot Audio-Visual Learning of Environment AcousticsSagnik Majumder, Changan Chen, Ziad Al-Halah, Kristen GraumanNeurIPS 2022 · 80 citations
- Mix and Localize: Localizing Sound Sources in MixturesXixi Hu, Ziyang Chen, Andrew OwensCVPR 2022 · 50 citations
- Modality-Independent Teachers Meet Weakly-Supervised Audio-Visual Event ParserYung-Hsuan Lai, Yen-Chun Chen, Frank WangNeurIPS 2023 · 27 citations
- Sound Localization from Motion: Jointly Learning Sound Direction and Camera RotationZiyang Chen, Shengyi Qian, Andrew OwensICCV 2023 · 21 citations
- Revisit Weakly-Supervised Audio-Visual Video Parsing from the Language PerspectiveYingying Fan, Yu Wu, Bo Du, Yutian LinNeurIPS 2023 · 20 citations
Builds on14
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- Object Goal Navigation using Goal-Oriented Semantic ExplorationDevendra Singh Chaplot, Dhiraj Gandhi, Abhinav Gupta, Ruslan SalakhutdinovNeurIPS 2020 · 857 citations
- DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion FramesErik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee et al.ICLR 2020 · 608 citations
- Co-Separating Sounds of Visual ObjectsRuohan Gao, Kristen GraumanICCV 2019 · 224 citations
- Bayesian Relational Memory for Semantic Visual NavigationYi Wu, Yuxin Wu, Aviv Tamar, Stuart Russell et al.ICCV 2019 · 114 citations
Related papers
- Learning to Set Waypoints for Audio-Visual NavigationChangan Chen, Sagnik Majumder, Ziad Al-Halah, Ruohan Gao et al.ICLR 2021 · 28 citations
- MARS-Sep: Multimodal-Aligned Reinforced Sound SeparationZihan Zhang, Xize Cheng, Zhennan Jiang, Dongjie Fu et al.ICLR 2026 · 2 citations
- RL-L: A Deep Reinforcement Learning Approach Intended for AR Label Placement in Dynamic ScenariosChen Zhu-Tian, Daniele Chiappalupi, Tica Lin, Yalong Yang et al.IEEE VIS 2023 · 14 citations
- Chat2Map: Efficient Scene Mapping from Multi-Ego ConversationsSagnik Majumder, Hao Jiang, Pierre Moulon, Ethan Henderson et al.CVPR 2023
- Learning Active Camera for Multi-Object NavigationPeihao Chen, Dongyu Ji, Kunyang Lin, Weiwen Hu et al.NeurIPS 2022 · 40 citations
