EE-RL: Vision Language Guided Reinforcement Learning with Explorer and Expert model for End-to-End Autonomous Driving
Xiaolong Li, Lan Yang, Ruyang Li, Shan Fang, Yang Liu, Xiangmo Zhao
Abstract
End-to-end driving frameworks, which directly map raw sensor data to vehicle control commands, have shown remarkable potential. However, their performance often deteriorates in sparse-critical scenarios, where rare but safetysensitive events occur. To address this problem, we propose Explorer-Expert Reinforcement Learning (EE-RL), a novel end-to-end framework that integrates an RL-based explorer, a fine-tuned vision-language model (VLM)-based expert, and a dual replay buffer. EE-RL adopts a collaborative learning strategy in which an explorer and two experts jointly generate experiences from regular driving scenarios to guide policy learning. As training progresses, one of the VLMs focuses on reasoning about sparse-critical scenarios, thereby enhancing learning efficiency and policy optimization in both scenarios. Additionally, the StateHash algorithm is designed to separately measure RGB image similarity and vehicle kinematic similarity, thereby skipping unnecessary VLM reasoning and enabling denser, more effective expert experience generation. Experiments on the CARLA Leaderboard demonstrate that EE-RL significantly outperforms state-of-the-art (SOTA) baselines, achieving +19.82% and +20.98% improvements in driving and infraction scores on Town03, respectively. The EE-RL further achieves 0% accident probability for red-light running and an average driving score of 80.09 in Town05-06, demonstrating robustness, generalizability, and the capability to handle sparse-critical scenarios. The code repository can be found in https://github.com/CAVTestLab/ EE-RL
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6dc39bde-3a0a-45b8-b3c2-35dff59ee674Builds on21
- VAD: Vectorized Scene Representation for Efficient Autonomous DrivingBo Jiang, Shaoyu Chen, Qing Xu, Bencheng Liao et al.ICCV 2023 · 602 citations
- Exploring Object-Centric Temporal Modeling for Efficient Multi-View 3D Object DetectionShihao Wang, Yingfei Liu, Tiancai Wang, Ying Li et al.ICCV 2023 · 399 citations
- End-to-End Urban Driving by Imitating a Reinforcement Learning CoachZhejun Zhang, Alexander Liniger, Dengxin Dai, Fisher Yu et al.ICCV 2021 · 313 citations
- Learning to drive from a world on railsDian Chen, Vladlen Koltun, Philipp KrähenbühlICCV 2021 · 164 citations
- Vision-Language Models are Zero-Shot Reward Models for Reinforcement LearningJuan Rocamonde, Victoriano Montesinos, Elvis Nava, Ethan Perez et al.ICLR 2024 · 154 citations
Related papers
- Exploring Data Aggregation in Policy Learning for Vision-Based Urban Autonomous DrivingAditya Prakash, Aseem Behl, Eshed Ohn-Bar, Kashyap Chitta et al.CVPR 2020
- Raw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2)Zhenjie Yang, Xiaosong Jia, Qifeng Li, Xue Yang et al.NeurIPS 2025 · 65 citations
- COVR: Collaborative Optimization of VLMs and RL Agent for Visual-Based ControlCanming Xia, Peixi Peng, Guang Tan, Zhan Su et al.AAAI 2026
- Unleashing VLA Potentials in Autonomous Driving via Explicit Learning from FailuresYuechen Luo, Fang Li, Qimao Chen, Shaoqing Xu et al.CVPR 2026 · 21 citations
- AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-TuningZewei Zhou, Tianhui Cai, Seth Z. Zhao, Yun Zhang et al.NeurIPS 2025 · 310 citations
