CogNav: Cognitive Process Modeling for Object Goal Navigation with LLMs
Yihan Cao, Jiazhao Zhang, Zhinan Yu, Shuzhen Liu, Zheng Qin, Qin Zou, Bo Du, Kai Xu
摘要
Object goal navigation (ObjectNav) is a fundamental task in embodied AI, requiring an agent to locate a target object in previously unseen environments. This task is particularly challenging because it requires both perceptual and cognitive processes, including object recognition and decision-making. While substantial advancements in perception have been driven by the rapid development of visual foundation models, progress on the cognitive aspect remains constrained, primarily limited to either implicit learning through simulator rollouts or explicit reliance on predefined heuristic rules. Inspired by neuroscientific findings demonstrating that humans maintain and dynamically update fine-grained cognitive states during object search tasks in novel environments, we propose CogNav, a framework designed to mimic this cognitive process using large language models. Specifically, we model the cognitive process using a finite state machine comprising fine-grained cognitive states, ranging from exploration to identification. Transitions between states are determined by a large language model based on a dynamically constructed heterogeneous cognitive map, which contains spatial and semantic information about the scene being explored. Extensive evaluations on the HM3D, MP3D, and RoboTHOR benchmarks demonstrate that our cognitive process modeling significantly improves the success rate of ObjectNav at least by relative 14% over the state-of-the-arts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- OpenFly: A COMPREHENSIVE PLATFORM FOR AERIAL VISION-LANGUAGE NAVIGATIONYunpeng Gao, Chenhui Li, Zhongrui You, Junli Liu 等ICLR 2026 · 被引用 61 次
- Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied ExplorationSen Wang, Bangwei Liu, Zhenkun Gao, Lizhuang Ma 等CVPR 2026 · 被引用 14 次
- CompassNav: Steering From Path Imitation to Decision Understanding In NavigationLinfeng Li, Jian Zhao, Yuan Xie, Xin Tan 等ICLR 2026 · 被引用 14 次
- CitySeeker: How Do VLMs Explore Embodied Urban Navigation with Implicit Human Needs?Siqi Wang, Chao Liang, Yunfan Gao, Erxin Yu 等ICLR 2026 · 被引用 8 次
- EfficientNav: Towards On-Device Object-Goal Navigation with Navigation Map Caching and RetrievalZebin Yang, Sunjian Zheng, Tong Xie, Tianshi Xu 等NeurIPS 2025 · 被引用 7 次
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Object Goal Navigation using Goal-Oriented Semantic ExplorationDevendra Singh Chaplot, Dhiraj Gandhi, Abhinav Gupta, Ruslan SalakhutdinovNeurIPS 2020 · 被引用 857 次
- Habitat 2.0: Training Home Assistants to Rearrange their HabitatAndrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans 等NeurIPS 2021 · 被引用 826 次
相关 Paper
- Distilling LLM Prior to Flow Model for Generalizable Agent's Imagination in Object Goal NavigationBadi Li, Renjie Lu, Yu Zhou, Jingke Meng 等NeurIPS 2025 · 被引用 5 次
- RoboTron-Nav: A Unified Framework for Embodied Navigation Integrating Perception, Planning, and PredictionYufeng Zhong, Chengjian Feng, Feng Yan, Fanfan Liu 等ICCV 2025 · 被引用 1 次
- BeliefMapNav: 3D Voxel-Based Belief Map for Zero-Shot Object NavigationZibo Zhou, Yue Hu, Lingkai Zhang, Zonglin Li 等NeurIPS 2025 · 被引用 31 次
- TANGO: Training-free Embodied AI Agents for Open-world TasksFilippo Ziliotto, Tommaso Campari, Luciano Serafini, Lamberto BallanCVPR 2025
- Embodied Navigation Foundation ModelJiazhao Zhang, Anqi Li, Yunpeng Qi, Minghan Li 等ICLR 2026 · 被引用 93 次
