See, Hear, Explore: Curiosity via Audio-Visual Association
Victoria Dean, Shubham Tulsiani, Abhinav Gupta
Abstract
Exploration is one of the core challenges in reinforcement learning. A common formulation of curiosity-driven exploration uses the difference between the real future and the future predicted by a learned model [1] . However, predicting the future is an inherently difficult task which can be ill-posed in the face of stochasticity. In this paper, we introduce an alternative form of curiosity that rewards novel associations between different senses. Our approach exploits multiple modalities to provide a stronger signal for more efficient exploration. Our method is inspired by the fact that, for humans, both sight and sound play a critical role in exploration. We present results on several Atari environments and Habitat (a photorealistic navigation simulator), showing the benefits of using an audio-visual association model for intrinsically guiding learning agents in the absence of external rewards. For videos and code, see https://vdean.github.io/audio-curiosity.html .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1d81fc01-6b93-4b01-b14c-8a2967bb6eecCited by top-tier papers11
- Few-Shot Audio-Visual Learning of Environment AcousticsSagnik Majumder, Changan Chen, Ziad Al-Halah, Kristen GraumanNeurIPS 2022 · 80 citations
- Toward Practical Monocular Indoor Depth EstimationCho-Ying Wu, Jialiang Wang, Michael Hall, Ulrich Neumann et al.CVPR 2022 · 68 citations
- Sound Adversarial Audio-Visual NavigationYinfeng Yu, Wenbing Huang, Fuchun Sun, Changan Chen et al.ICLR 2022 · 49 citations
- Visual Acoustic MatchingChangan Chen, Ruohan Gao, Paul Calamia, Kristen GraumanCVPR 2022 · 42 citations
- Learning Active Camera for Multi-Object NavigationPeihao Chen, Dongyu Ji, Kunyang Lin, Weiwen Hu et al.NeurIPS 2022 · 40 citations
Builds on1
Related papers
- SENSEI: Semantic Exploration Guided by Foundation Models to Learn Versatile World ModelsCansu Sancaktar, Christian Gumbsch, Andrii Zadaianchuk, Pavel Kolev et al.ICML 2025
- GeoExplorer: Active Geo-Localization with Curiosity-Driven ExplorationLi Mi, Manon Béchaz, Zeming Chen, Antoine Bosselut et al.ICCV 2025
- What You Think is What You See: Driving Exploration in VLM Agents via Visual-Linguistic CuriosityHaoxi Li, Qinglin Hou, Jianfei Ma, Jinxiang Lai et al.ICML 2026
- Learning to Set Waypoints for Audio-Visual NavigationChangan Chen, Sagnik Majumder, Ziad Al-Halah, Ruohan Gao et al.ICLR 2021 · 28 citations
- How to Stay Curious while avoiding Noisy TVs using Aleatoric Uncertainty EstimationAugustine N. Mavor-Parker, Kimberly A. Young, Caswell Barry, Lewis D. GriffinICML 2022 · 32 citations
