LASAR: Towards Spatio-temporal Reasoning with Latent Cognitive Map
Jinzhou Tang, Sidi Liu, Waikit Xiu, Weixing Chen, Keze Wang
摘要
A fundamental challenge in embodied AI is verifying if agents build internal models of spatial structure or merely learn to mimic task-specific expert trajectories. This is critical as foundational approaches rooted in action-centric tasks (e.g., VLN) and reasoning-centric tasks (e.g., EQA) often share a common limitation: they lack a learning signal that forces them to encode fine-grained spatial relationships (like topology or distance) over long-range, fragmented experiences. To address this, we first propose LASAR, an architecture featuring a dual-memory system designed to maintain both episodic experiences and a semantic cognitive map. We then introduce Spatio-temporal Contextual Representation Learning (ST-CRL), a contrastive objective designed to train this architecture. ST-CRL leverages spatio-temporal cues from cognitive queries generated through annotated spatio-temporal context in simulation to build sample pairs, thereby forming the internal cognitive map from the agent's experiences. Experiments demonstrate that our method achieves 2%-3.5% gains in both zero-shot generalization on standard VLN-CE and VSI-Bench benchmarks. We also demonstrate that our proposed cognitive map has high self-consistency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
- EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of ThoughtYao Mu, Qinglong Zhang, Mengkang Hu, Wenhai Wang 等NeurIPS 2023 · 被引用 453 次
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language ModelsGengze Zhou, Yicong Hong, Qi WuAAAI 2024 · 被引用 361 次
- Spatio-temporal Self-Supervised Representation Learning for 3D Point CloudsSiyuan Huang, Yichen Xie, Song-Chun Zhu, Yixin ZhuICCV 2021 · 被引用 259 次
相关 Paper
- Learning Navigational Visual Representations with Semantic Map SupervisionYicong Hong, Yang Zhou, Ruiyi Zhang, Franck Dernoncourt 等ICCV 2023 · 被引用 56 次
- Semantic MapNet: Building Allocentric Semantic Maps and Representations from Egocentric ViewsVincent Cartillier, Zhile Ren, Neha Jain, Stefan Lee 等AAAI 2021 · 被引用 89 次
- Multi-Scale Gaussian-Language Map for Zero-shot Embodied Navigation and ReasoningSixian Zhang, Yiyao Wang, Xinhang Song, Keming Zhang 等CVPR 2026 · 被引用 3 次
- MapNav: A Novel Memory Representation via Annotated Semantic Maps for VLM-based Vision-and-Language NavigationLingfeng Zhang, Xiaoshuai Hao, Qinwen Xu, Qiang Zhang 等ACL 2025 · 被引用 55 次
- Temporal Representations for Exploration: Learning Complex Exploratory Behavior without Extrinsic RewardsFaisal Mohamed, Catherine Ji, Benjamin Eysenbach, Glen BersethICLR 2026 · 被引用 1 次
