RLfOLD: Reinforcement Learning from Online Demonstrations in Urban Autonomous Driving
Daniel Coelho, Miguel Oliveira, Vitor Santos
摘要
Reinforcement Learning from Demonstrations (RLfD) has emerged as an effective method by fusing expert demonstrations into Reinforcement Learning (RL) training, harnessing the strengths of both Imitation Learning (IL) and RL. However, existing algorithms rely on offline demonstrations, which can introduce a distribution gap between the demonstrations and the actual training environment, limiting their performance. In this paper, we propose a novel approach, Reinforcement Learning from Online Demonstrations (RL-fOLD), that leverages online demonstrations to address this limitation, ensuring the agent learns from relevant and up-todate scenarios, thus effectively bridging the distribution gap. Unlike conventional policy networks used in typical actorcritic algorithms, RLfOLD introduces a policy network that outputs two standard deviations: one for exploration and the other for IL training. This novel design allows the agent to adapt to varying levels of uncertainty inherent in both RL and IL. Furthermore, we introduce an exploration process guided by an online expert, incorporating an uncertainty-based technique. Our experiments on the CARLA NoCrash benchmark demonstrate the effectiveness and efficiency of RLfOLD. Notably, even with a significantly smaller encoder and a singlecamera setup, RLfOLD surpasses state-of-the-art methods in this evaluation. These results, achieved with limited resources, highlight RLfOLD as a highly promising solution for real-world applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Diffusion Imitation from ObservationBo-Ruei Huang, Chun-Kai Yang, Chun-Mao Lai, Dai-Jie Wu 等NeurIPS 2024 · 被引用 15 次
- LPPG-RL: Lexicographically Projected Policy Gradient Reinforcement Learning with Subproblem ExplorationRuiyu Qiu, Rui Wang, Guanghui Yang, Xiang Li 等AAAI 2026
- Master Skill Learning with Policy-Grounded Synergy of LLM-based Reward Shaping and ExploringYanbin Chang, Junfan Lin, Jie Jiang, Runhao Zeng 等ICLR 2026
它引用的顶会 Paper12
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 被引用 911 次
- Exploring the Limitations of Behavior Cloning for Autonomous DrivingFelipe Codevilla, Eder Santana, Antonio M. López, Adrien GaidonICCV 2019 · 被引用 666 次
- Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement LearningDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICLR 2022 · 被引用 457 次
- End-to-End Urban Driving by Imitating a Reinforcement Learning CoachZhejun Zhang, Alexander Liniger, Dengxin Dai, Fisher Yu 等ICCV 2021 · 被引用 313 次
- Learning to drive from a world on railsDian Chen, Vladlen Koltun, Philipp KrähenbühlICCV 2021 · 被引用 164 次
相关 Paper
- Reinforcement Learning from Imperfect Demonstrations under Soft Expert GuidanceMingxuan Jing, Xiaojian Ma, Wenbing Huang, Fuchun Sun 等AAAI 2020 · 被引用 70 次
- Learning and Repair of Deep Reinforcement Learning Policies from Fuzz-Testing DataMartin Tappler, Andrea Pferscher, Bernhard K. Aichernig, Bettina KönighoferICSE 2024 · 被引用 6 次
- Hybrid Policy Optimization from Imperfect DemonstrationsHanlin Yang, Chao Yu, Peng Sun, Siji ChenNeurIPS 2023 · 被引用 14 次
- Uncertainty-Guided Exploration and Stable Planning for Sparse-Reward Manipulation from Limited DemonstrationsHaowen Sun, Liqi Huang, Mingyang Li, Sihua Ren 等ICML 2026
- Uncertainty-Based Smooth Policy Regularisation for Reinforcement Learning with Few DemonstrationsYujie Zhu, Charles A. Hepburn, Matthew Thorpe, Giovanni MontanaNeurIPS 2025
