Environment Predictive Coding for Visual Navigation
Santhosh Kumar Ramakrishnan, Tushar Nagarajan, Ziad Al-Halah, Kristen Grauman
Abstract
We introduce environment predictive coding, a self-supervised approach to learn environment-level representations for embodied agents. In contrast to prior work on self-supervised learning for individual images, we aim to encode a 3D environment using a series of images observed by an agent moving in it. We learn these representations via a masked-zone prediction task, which segments an agent’s trajectory into zones and then predicts features of randomly masked zones, conditioned on the agent’s camera poses. This explicit spatial conditioning encourages learning representations that capture the geometric and semantic regularities of 3D environments. We learn such representations on a collection of video walkthroughs and demonstrate successful transfer to multiple downstream navigation tasks. Our experiments on the real-world scanned 3D environments of Gibson and Matterport3D show that our method obtains 2 - 6× higher sample-efficiency and up to 57% higher performance over standard image-representation learning.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 3c864288-407f-4774-a95b-b3f0031b7a46Cited by top-tier papers4
- Multi-label affordance mapping from egocentric visionLorenzo Mur-Labadia, Josechu J. Guerrero, Ruben Martinez-CantinICCV 2023 · 26 citations
- ENTL: Embodied Navigation Trajectory LearnerKlemen Kotar, Aaron Walsman, Roozbeh MottaghiICCV 2023 · 16 citations
- DIV-FF: Dynamic Image-Video Feature Fields For Environment Understanding in Egocentric VideosLorenzo Mur-Labadia, Josechu Guerrero, Ruben Martinez-CantinCVPR 2025
- MemoNav: Working Memory Model for Visual NavigationHongxin Li, Zeyu Wang, Xu Yang, Yuran Yang et al.CVPR 2024
Related papers
- Imagine Before Go: Self-Supervised Generative Map for Object Goal NavigationSixian Zhang, Xinyao Yu, Xinhang Song, Xiaohan Wang et al.CVPR 2024 · 14 citations
- Embodied Contrastive Learning with Geometric Consistency and Behavioral Awareness for Object NavigationBolei Chen, Jiaxu Kang, Ping Zhong, Yixiong Liang et al.ACM MM 2024 · 4 citations
- Scene Graph Contrastive Learning for Embodied NavigationKunal Pratap Singh, Jordi Salvador, Luca Weihs, Aniruddha KembhaviICCV 2023 · 31 citations
- Embodied Visual Active Learning for Semantic SegmentationDavid Nilsson, Aleksis Pirinen, Erik Gärtner, Cristian SminchisescuAAAI 2021 · 37 citations
- Learning Navigational Visual Representations with Semantic Map SupervisionYicong Hong, Yang Zhou, Ruiyi Zhang, Franck Dernoncourt et al.ICCV 2023 · 56 citations
