EE-RL: Vision Language Guided Reinforcement Learning with Explorer and Expert model for End-to-End Autonomous Driving
Xiaolong Li, Lan Yang, Ruyang Li, Shan Fang, Yang Liu, Xiangmo Zhao
摘要
End-to-end driving frameworks, which directly map raw sensor data to vehicle control commands, have shown remarkable potential. However, their performance often deteriorates in sparse-critical scenarios, where rare but safetysensitive events occur. To address this problem, we propose Explorer-Expert Reinforcement Learning (EE-RL), a novel end-to-end framework that integrates an RL-based explorer, a fine-tuned vision-language model (VLM)-based expert, and a dual replay buffer. EE-RL adopts a collaborative learning strategy in which an explorer and two experts jointly generate experiences from regular driving scenarios to guide policy learning. As training progresses, one of the VLMs focuses on reasoning about sparse-critical scenarios, thereby enhancing learning efficiency and policy optimization in both scenarios. Additionally, the StateHash algorithm is designed to separately measure RGB image similarity and vehicle kinematic similarity, thereby skipping unnecessary VLM reasoning and enabling denser, more effective expert experience generation. Experiments on the CARLA Leaderboard demonstrate that EE-RL significantly outperforms state-of-the-art (SOTA) baselines, achieving +19.82% and +20.98% improvements in driving and infraction scores on Town03, respectively. The EE-RL further achieves 0% accident probability for red-light running and an average driving score of 80.09 in Town05-06, demonstrating robustness, generalizability, and the capability to handle sparse-critical scenarios. The code repository can be found in https://github.com/CAVTestLab/ EE-RL
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- VAD: Vectorized Scene Representation for Efficient Autonomous DrivingBo Jiang, Shaoyu Chen, Qing Xu, Bencheng Liao 等ICCV 2023 · 被引用 602 次
- Exploring Object-Centric Temporal Modeling for Efficient Multi-View 3D Object DetectionShihao Wang, Yingfei Liu, Tiancai Wang, Ying Li 等ICCV 2023 · 被引用 399 次
- End-to-End Urban Driving by Imitating a Reinforcement Learning CoachZhejun Zhang, Alexander Liniger, Dengxin Dai, Fisher Yu 等ICCV 2021 · 被引用 313 次
- Learning to drive from a world on railsDian Chen, Vladlen Koltun, Philipp KrähenbühlICCV 2021 · 被引用 164 次
- Vision-Language Models are Zero-Shot Reward Models for Reinforcement LearningJuan Rocamonde, Victoriano Montesinos, Elvis Nava, Ethan Perez 等ICLR 2024 · 被引用 154 次
相关 Paper
- Exploring Data Aggregation in Policy Learning for Vision-Based Urban Autonomous DrivingAditya Prakash, Aseem Behl, Eshed Ohn-Bar, Kashyap Chitta 等CVPR 2020
- Raw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2)Zhenjie Yang, Xiaosong Jia, Qifeng Li, Xue Yang 等NeurIPS 2025 · 被引用 65 次
- COVR: Collaborative Optimization of VLMs and RL Agent for Visual-Based ControlCanming Xia, Peixi Peng, Guang Tan, Zhan Su 等AAAI 2026
- Unleashing VLA Potentials in Autonomous Driving via Explicit Learning from FailuresYuechen Luo, Fang Li, Qimao Chen, Shaoqing Xu 等CVPR 2026 · 被引用 21 次
- AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-TuningZewei Zhou, Tianhui Cai, Seth Z. Zhao, Yun Zhang 等NeurIPS 2025 · 被引用 310 次
