Efficient Learning of Safe Driving Policy via Human-AI Copilot Optimization
Quanyi Li, Zhenghao Peng, Bolei Zhou
摘要
Human intervention is an effective way to inject human knowledge into the loop of reinforcement learning, bringing fast learning and training safety. But given the very limited budget of human intervention, it is challenging to design when and how human expert interacts with the learning agent in the training. In this work, we develop a novel human-in-the-loop learning method called Human-AI Copilot Optimization (HACO). To allow the agent's sufficient exploration in the risky environments while ensuring the training safety, the human expert can take over the control and demonstrate to the agent how to avoid probably dangerous situations or trivial behaviors. The proposed HACO then effectively utilizes the data collected both from the trial-and-error exploration and human's partial demonstration to train a high-performing agent. HACO extracts proxy state-action values from partial human demonstration and optimizes the agent to improve the proxy values while reducing the human interventions. No environmental reward is required in HACO. The experiments show that HACO achieves a substantially high sample efficiency in the safe driving benchmark. It can train agents to drive in unseen traffic scenes with a handful of human intervention budget and achieve high safety and generalizability, outperforming both reinforcement learning and imitation learning baselines with a large margin. Code and demo videos are available at: https://decisionforce.github.io/HACO/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- RLIF: Interactive Imitation Learning as Reinforcement LearningJianlan Luo, Perry Dong, Yuexiang Zhai, Yi Ma 等ICLR 2024 · 被引用 31 次
- Learning from Active Human Involvement through Proxy Value PropagationZhenghao Mark Peng, Wenjie Mo, Chenda Duan, Quanyi Li 等NeurIPS 2023 · 被引用 30 次
- Human-assisted Robotic Policy Refinement via Action Preference OptimizationWenke Xia, Yichu Yang, Hongtao Wu, Xiao Ma 等NeurIPS 2025 · 被引用 17 次
- RLfOLD: Reinforcement Learning from Online Demonstrations in Urban Autonomous DrivingDaniel Coelho, Miguel Oliveira, Vitor SantosAAAI 2024 · 被引用 14 次
- Shared Autonomy with IDA: Interventional Diffusion AssistanceBrandon McMahan, Zhenghao Mark Peng, Bolei Zhou, Jonathan C. KaoNeurIPS 2024 · 被引用 12 次
它引用的顶会 Paper6
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Responsive Safety in Reinforcement Learning by PID Lagrangian MethodsAdam Stooke, Joshua Achiam, Pieter AbbeelICML 2020 · 被引用 403 次
- SQIL: Imitation Learning via Reinforcement Learning with Sparse RewardsSiddharth Reddy, Anca D. Dragan, Sergey LevineICLR 2020 · 被引用 299 次
- Widening the Pipeline in Human-Guided Reinforcement Learning with Explanation and Context-Aware Data AugmentationLin Guan, Mudit Verma, Sihang Guo, Ruohan Zhang 等NeurIPS 2021 · 被引用 57 次
- Continual Learning of Control Primitives : Skill Discovery via Reset-GamesKelvin Xu, Siddharth Verma, Chelsea Finn, Sergey LevineNeurIPS 2020 · 被引用 37 次
相关 Paper
- Robot-Gated Interactive Imitation Learning with Adaptive Intervention MechanismHaoyuan Cai, Zhenghao Peng, Bolei ZhouICML 2025
- Faithful Dynamic Imitation Learning from Human Intervention with Dynamic Regret MinimizationBo Ling, Zhengyu Gan, Wanyuan Wang, Guanyu Gao 等NeurIPS 2025
- Reinforcement Learning from Imperfect Corrective Actions and Proxy RewardsZhaohui Jiang, Xuening Feng, Paul Weng, Yifei Zhu 等ICLR 2025
- Predictive Preference Learning from Human InterventionsHaoyuan Cai, Zhenghao Mark Peng, Bolei ZhouNeurIPS 2025 · 被引用 6 次
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 被引用 380 次
