Embodied Contrastive Learning with Geometric Consistency and Behavioral Awareness for Object Navigation
Bolei Chen, Jiaxu Kang, Ping Zhong, Yixiong Liang, Yu Sheng, Jianxin Wang
Abstract
Object Navigation (ObjcetNav), which enables an agent to seek any instance of an object category specified by a semantic label, has shown great advances. However, current agents are built upon occlusion-prone visual observations or compressed 2D semantic maps, which hinder their embodied perception of 3D scene geometry and easily lead to ambiguous object localization and blind exploration. To address these limitations, we present an Embodied Contrastive Learning (ECL) method with Geometric Consistency (GC) and Behavioral Awareness (BA), which motivates agents to actively encode 3D scene layouts and semantic cues. Driven by our embodied exploration strategy, BA is modeled by predicting navigational actions based on multi-frame visual images, as behaviors that cause differences between adjacent visual sensations are crucial for learning correlations among continuous visions. The GC is modeled as the alignment of behavior-aware visual stimulus with 3D semantic shapes by employing unsupervised contrastive learning. The aligned behavior-aware visual features and geometric invariance priors are injected into a modular ObjectNav framework to enhance object recognition and exploration capabilities. As expected, our ECL method performs well on object detection and instance segmentation tasks. Our ObjectNav strategy outperforms state-of-the-art methods on MP3D and Gibson datasets, showing the potential of our ECL in embodied navigation.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 01773f3d-b070-4d93-a856-9af3f3cad789Cited by top-tier papers5
- Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement LearningBaining Zhao, Ziyou Wang, Jianjie Fang, Chen Gao et al.ACM MM 2025 · 8 citations
- SeqWalker: Sequential-Horizon Vision-and-Language Navigation with Hierarchical PlanningZebin Han, Xudong Wang, Baichen Liu, Qi Lyu et al.AAAI 2026 · 2 citations
- Continuous Vision-Language-Action Co-Learning with Semantic-Physical Alignment for Behavioral CloningXiuxiu Qi, Yu Yang, Jiannong Cao, Luyao Bai et al.AAAI 2026 · 2 citations
- Perspective from a Broader Context: Can Room Style Knowledge Help Visual Floorplan Localization?Bolei Chen, Shengsheng Yan, Yongzheng Cui, Jiaxu Kang et al.AAAI 2026 · 1 citation
- Perspective from a Higher Dimension: Can 3D Geometric Priors Help Visual Floorplan Localization?Bolei Chen, Jiaxu Kang, Haonan Yang, Ping Zhong et al.ACM MM 2025
Related papers
- 3D-Aware Object Goal Navigation via Simultaneous Exploration and IdentificationJiazhao Zhang, Liu Dai, Fanpeng Meng, Qingnan Fan et al.CVPR 2023
- Imagine Before Go: Self-Supervised Generative Map for Object Goal NavigationSixian Zhang, Xinyao Yu, Xinhang Song, Xiaohan Wang et al.CVPR 2024 · 14 citations
- Environment Predictive Coding for Visual NavigationSanthosh Kumar Ramakrishnan, Tushar Nagarajan, Ziad Al-Halah, Kristen GraumanICLR 2022 · 10 citations
- Scene Graph Contrastive Learning for Embodied NavigationKunal Pratap Singh, Jordi Salvador, Luca Weihs, Aniruddha KembhaviICCV 2023 · 31 citations
- Self-Supervised Image Representation Learning with Geometric Set ConsistencyNenglun Chen, Lei Chu, Hao Pan, Yan Lu et al.CVPR 2022 · 8 citations
