NaviFormer: A Spatio-Temporal Context-Aware Transformer for Object Navigation
Wei Xie, Haobo Jiang, Yun Zhu, Jianjun Qian, Jin Xie
Abstract
Learning discriminative state representations of agents, encompassing the spatial layout and temporal pose trajectory, is essential for effective navigation decisions. However, existing approaches often rely on simplistic plain networks for navigation information fusion, overlooking the complex long-range dependencies across spatio-temporal cues, which leads to suboptimal state perception and potential decision failures. In this paper, we introduce NaviFormer, an effective encoder-decoder navigation transformer, to aggregate discriminative spatio-temporal context information for object navigation. Our navigation encoder not only encodes spatial layouts and temporal agent poses but also innovatively constructs and encodes a passable frontier map, enriching the original state encoding with cues of potential exploration regions. Furthermore, our navigation decoder employs spatio-temporal self-attention and cross-attention mechanisms to model the dependencies among spatial layout encoding, temporal pose encoding, and passable frontier encoding, thereby facilitating comprehensive contextual state feature aggregation. Finally, we leverage these learned spatio-temporal contextual state representations for PPO-based navigation decisions. Extensive experiments on the Gibson, Habitat-Matterport3D (HM3D) and Matterport3D (MP3D) datasets demonstrate the superiority of our approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Few-Shot Incremental 3D Object Detection in Dynamic Indoor EnvironmentsYun Zhu, Jianjun Qian, Jian Yang, Jin Xie et al.CVPR 2026 · 2 citations
- Fisher-Preserving Guidance: Training-Free Manifold Constraints for Safe Diffusion ControlHao Ren, Zetong Bi, Yiming Zeng, Le Zheng et al.ICML 2026 · 1 citation
- Expand Your SCOPE: Semantic Cognition over Potential-Based Exploration for Embodied Visual NavigationNingnan Wang, Weihuang Chen, Liming Chen, Haoxuan Ji et al.AAAI 2026
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- Object Goal Navigation using Goal-Oriented Semantic ExplorationDevendra Singh Chaplot, Dhiraj Gandhi, Abhinav Gupta, Ruslan SalakhutdinovNeurIPS 2020 · 857 citations
- DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion FramesErik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee et al.ICLR 2020 · 608 citations
- Learning To Explore Using Active Neural SLAMDevendra Singh Chaplot, Dhiraj Gandhi, Saurabh Gupta, Abhinav Gupta et al.ICLR 2020 · 603 citations
Related papers
- Learning Navigational Visual Representations with Semantic Map SupervisionYicong Hong, Yang Zhou, Ruiyi Zhang, Franck Dernoncourt et al.ICCV 2023 · 56 citations
- PEANUT: Predicting and Navigating to Unseen TargetsAlbert J. Zhai, Shenlong WangICCV 2023 · 52 citations
- MemoNav: Working Memory Model for Visual NavigationHongxin Li, Zeyu Wang, Xu Yang, Yuran Yang et al.CVPR 2024
- KTPFormer: Kinematics and Trajectory Prior Knowledge-Enhanced Transformer for 3D Human Pose EstimationJihua Peng, Yanghong Zhou, P. Y. MokCVPR 2024 · 67 citations
- VTNet: Visual Transformer Network for Object Goal NavigationHeming Du, Xin Yu, Liang ZhengICLR 2021 · 36 citations
