ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving
Yongkang Li, Kaixin Xiong, Xiangyu Guo, Fang Li, Sixu Yan, Gangwei Xu, Lijun Zhou, Long Chen, Haiyang Sun, BING WANG, Kun Ma, Guang Chen
Abstract
Recent studies have explored leveraging the world knowledge and cognitive capabilities of Vision-Language Models (VLMs) to address the long-tail problem in end-to-end autonomous driving. However, existing methods typically formulate trajectory planning as a language modeling task, where physical actions are output in the language space, potentially leading to issues such as format-violating outputs, infeasible actions, and slow inference speeds. In this paper, we propose ReCogDrive, a novel Reinforced Cognitive framework for end-to-end autonomous Driving, unifying driving understanding and planning by integrating an autoregressive model with a diffusion planner. First, to instill human driving cognition into the VLM, we introduce a hierarchical data pipeline that mimics the sequential cognitive process of human drivers through three stages: generation, refinement, and quality control. Building on this cognitive foundation, we then address the language-action mismatch by injecting the VLM's learned driving priors into a diffusion planner to efficiently generate continuous and stable trajectories. Furthermore, to enhance driving safety and reduce collisions, we introduce a Diffusion Group Relative Policy Optimization (DiffGRPO) stage, reinforcing the planner for enhanced safety and comfort. Extensive experiments on the NAVSIM and Bench2Drive benchmarks demonstrate that ReCogDrive achieves state-of-the-art performance. Additionally, qualitative results across diverse driving scenarios and DriveBench highlight the model's scene comprehension. Code and models are available at https://github.com/xiaomi-research/recogdrive . INTRODUCTION Autonomous driving, which aims to predict a smooth, comfortable, and collision-free trajectory for a vehicle, has seen significant advancements. Recent end-to-end autonomous driving systems (Jiang
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5b6458ec-ec63-4d37-bfe6-d756ea19d867Cited by top-tier papers16
- DriveLaW: Unifying Planning and Video Generation in a Latent Driving WorldTianze Xia, Yongkang Li, Lijun Zhou, Jingfeng Yao et al.CVPR 2026 · 58 citations
- DriveDPO: Policy Learning via Safety DPO For End-to-End Autonomous DrivingShuyao Shang, Yuntao Chen, Yuqi Wang, Yingyan Li et al.NeurIPS 2025 · 49 citations
- Driving on RegistersEllington Kirby, Alexandre Boulch, Yihong Xu, Yuan Yin et al.CVPR 2026 · 45 citations
- SimScale: Learning to Drive via Real-World Simulation at ScaleHaochen Tian, Tianyu Li, Haochen Liu, Jiazhi Yang et al.CVPR 2026 · 40 citations
- Unleashing VLA Potentials in Autonomous Driving via Explicit Learning from FailuresYuechen Luo, Fang Li, Qimao Chen, Shaoqing Xu et al.CVPR 2026 · 21 citations
Builds on32
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
Related papers
- SGDrive: Scene-to-Goal Hierarchical World Cognition for Autonomous Drivingjingyu li, Junjie Wu, Dongnan Hu, Xiangkai Huang et al.CVPR 2026 · 36 citations
- AutoDrive-P3: Unified Chain of Perception-Prediction-Planning Thought via Reinforcement Fine-TuningYuqi Ye, Zijian Zhang, Junhong Lin, Shangkun Sun et al.ICLR 2026 · 16 citations
- AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-TuningZewei Zhou, Tianhui Cai, Seth Z. Zhao, Yun Zhang et al.NeurIPS 2025 · 310 citations
- ReAL-AD: Towards Human-Like Reasoning in End-to-End Autonomous DrivingYuhang Lu, Jiadong Tu, Yuexin Ma, Xinge ZhuICCV 2025 · 1 citation
- Discrete Diffusion for Reflective Vision-Language-Action Models in Autonomous DrivingPengxiang Li, Yinan Zheng, Yue Wang, Huimin Wang et al.ICLR 2026 · 24 citations
