CogDriver: Integrating Cognitive Inertia for Temporally Coherent Planning in Autonomous Driving
Pei Liu, Qingtian Ning, Xinyan Lu, Haipeng Liu, Weiliang Ma, Dangen She, Xianpeng Lang, Jun Ma
摘要
The pursuit of autonomous agents capable of temporally coherent planning is hindered by a fundamental flaw in current vision-language models (VLMs): they lack cognitive inertia. Operating on isolated snapshots, these models cannot form a continuous understanding of the environment, leading to erratic decision jitter and a failure to execute complex, multi-step maneuvers. To remedy this, we introduce CogDriver, a framework designed to build a stable internal representation by instilling this crucial cognitive property. Our work makes two key contributions: (1) We present CogDriver-Data, a large-scale vision-language-action dataset whose narrative annotations provide the supervisory signal for learning temporal dynamics and persistent intent. (2) We develop the CogDriver-Agent, an architecture featuring a sparse temporal memory to maintain a stable internal state. This is enabled by a spatiotemporal knowledge distillation approach that explicitly teaches decision coherence. Comprehensive experiments validate our paradigm: CogDriver-Agent achieves a 22% increase in the closed-loop Driving Score on Bench2Drive and a 21% reduction in mean L2 error on nuScenes, establishing a new state-of-the-art. These significant gains in both long-term decision-making and imitation accuracy provide strong evidence that our agent successfully maintains a temporally coherent internal state, bridging the gap toward more reliable autonomous driving. Project link: https://ocean-luna.github.io/CogDriver.github.io/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- E3AD: An Emotion-Aware Vision-Language-Action Model for Human-Centric End-to-End Autonomous DrivingYihong Tang, Haicheng Liao, Tong Nie, Junlin He 等CVPR 2026 · 被引用 2 次
- CaT-GS: Efficient 3DGS Rendering for Large-Scale Scenes with Inter-frame Caching and Tile SchedulingTingjia Zhang, Bo Chen, Shengzhong Liu, Fan Wu 等CVPR 2026
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- VAD: Vectorized Scene Representation for Efficient Autonomous DrivingBo Jiang, Shaoyu Chen, Qing Xu, Bencheng Liao 等ICCV 2023 · 被引用 602 次
- Trajectory-guided Control Prediction for End-to-end Autonomous Driving: A Simple yet Strong BaselinePenghao Wu, Xiaosong Jia, Li Chen, Junchi Yan 等NeurIPS 2022 · 被引用 444 次
- VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic PlanningBo Jiang, Shaoyu Chen, Hao Gao, Bencheng Liao 等ICLR 2026 · 被引用 259 次
相关 Paper
- MindDriver: Introducing Progressive Multimodal Reasoning for Autonomous DrivingLingjun Zhang, Yujian Yuan, Changjie Wu, Xinyuan Chang 等CVPR 2026 · 被引用 13 次
- FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous DrivingShuang Zeng, Xinyuan Chang, Mengwei Xie, Xinran Liu 等NeurIPS 2025 · 被引用 228 次
- VLR-Driver: Large Vision-Language-Reasoning Models for Embodied Autonomous DrivingFanjie Kong, Yitong Li, Weihuang Chen, Chen Min 等ICCV 2025 · 被引用 2 次
- SGDrive: Scene-to-Goal Hierarchical World Cognition for Autonomous Drivingjingyu li, Junjie Wu, Dongnan Hu, Xiangkai Huang 等CVPR 2026 · 被引用 36 次
- ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous DrivingYongkang Li, Kaixin Xiong, Xiangyu Guo, Fang Li 等ICLR 2026 · 被引用 196 次
