Learning to Move with Affordance Maps
William Qi, Ravi Teja Mullapudi, Saurabh Gupta, Deva Ramanan
Abstract
The ability to autonomously explore and navigate a physical space is a fundamental requirement for virtually any mobile autonomous agent, from household robotic vacuums to autonomous vehicles. Traditional SLAM-based approaches for exploration and navigation largely focus on leveraging scene geometry, but fail to model dynamic objects (such as other agents) or semantic constraints (such as wet floors or doorways). Learning-based RL agents are an attractive alternative because they can incorporate both semantic and geometric information, but are notoriously sample inefficient, difficult to generalize to novel settings, and are difficult to interpret. In this paper, we combine the best of both worlds with a modular approach that learns a spatial representation of a scene that is trained to be effective when coupled with traditional geometric planners. Specifically, we design an agent that learns to predict a spatial affordance map that elucidates what parts of a scene are navigable through active self-supervised experience gathering. In contrast to most simulation environments that assume a static world, we evaluate our approach in the VizDoom simulator, using large-scale randomly-generated maps containing a variety of dynamic actors and hazards. We show that learned affordance maps can be used to augment traditional approaches for both exploration and navigation, providing significant improvements in performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Learning Affordance Landscapes for Interaction Exploration in 3D EnvironmentsTushar Nagarajan, Kristen GraumanNeurIPS 2020 · 87 citations
- SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object ManipulationZekun Qi, Wenyao Zhang, Yufei Ding, Runpei Dong et al.NeurIPS 2025 · 65 citations
- Affordances-Oriented Planning Using Foundation Models for Continuous Vision-Language NavigationJiaqi Chen, Bingqian Lin, Xinmin Liu, Lin Ma et al.AAAI 2025 · 61 citations
- Embodied Visual Active Learning for Semantic SegmentationDavid Nilsson, Aleksis Pirinen, Erik Gärtner, Cristian SminchisescuAAAI 2021 · 37 citations
- Shaping embodied agent behavior with activity-context priors from egocentric videoTushar Nagarajan, Kristen GraumanNeurIPS 2021 · 23 citations
Related papers
- An Interactive Navigation Method with Effect-oriented AffordanceXiaohan Wang, Yuehu Liu, Xinhang Song, Yuyi Liu et al.CVPR 2024 · 1 citation
- Multi-Object Navigation with dynamically learned neural implicit representationsPierre Marza, Laëtitia Matignon, Olivier Simonin, Christian WolfICCV 2023 · 32 citations
- Imagine Before Go: Self-Supervised Generative Map for Object Goal NavigationSixian Zhang, Xinyao Yu, Xinhang Song, Xiaohan Wang et al.CVPR 2024 · 14 citations
- Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language NavigationKailing Li, Tianwen Qian, Lijin Yang, Yuqian Fu et al.CVPR 2026 · 10 citations
- GAMap: Zero-Shot Object Goal Navigation with Multi-Scale Geometric-Affordance GuidanceShuaihang Yuan, Hao Huang, Yu Hao, Congcong Wen et al.NeurIPS 2024 · 42 citations
