Guarded Policy Optimization with Imperfect Online Demonstrations
Zhenghai Xue, Zhenghao Peng, Quanyi Li, Zhihan Liu, Bolei Zhou
摘要
The Teacher-Student Framework (TSF) is a reinforcement learning setting where a teacher agent guards the training of a student agent by intervening and providing online demonstrations. Assuming optimal, the teacher policy has the perfect timing and capability to intervene in the learning process of the student agent, providing safety guarantee and exploration guidance. Nevertheless, in many real-world settings it is expensive or even impossible to obtain a well-performing teacher policy. In this work, we relax the assumption of a well-performing teacher and develop a new method that can incorporate arbitrary teacher policies with modest or inferior performance. We instantiate an Off-Policy Reinforcement Learning algorithm, termed Teacher-Student Shared Control (TS2C), which incorporates teacher intervention based on trajectory-based value estimation. Theoretical analysis validates that the proposed TS2C algorithm attains efficient exploration and substantial safety guarantee without being affected by the teacher's own performance. Experiments on various continuous control tasks show that our method can exploit teacher policies at different performance levels while maintaining a low training cost. Moreover, the student policy surpasses the imperfect teacher policy in terms of higher accumulated reward in held-out testing environments. Code is available at https://metadriverse.github.io/TS2C .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- RLIF: Interactive Imitation Learning as Reinforcement LearningJianlan Luo, Perry Dong, Yuexiang Zhai, Yi Ma 等ICLR 2024 · 被引用 31 次
- Residual Q-Learning: Offline and Online Policy Customization without ValueChenran Li, Chen Tang, Haruki Nishimura, Jean Mercat 等NeurIPS 2023 · 被引用 15 次
- Shared Autonomy with IDA: Interventional Diffusion AssistanceBrandon McMahan, Zhenghao Mark Peng, Bolei Zhou, Jonathan C. KaoNeurIPS 2024 · 被引用 12 次
- Predictive Preference Learning from Human InterventionsHaoyuan Cai, Zhenghao Mark Peng, Bolei ZhouNeurIPS 2025 · 被引用 6 次
- Policy Regularization on Globally Accessible States in Cross-Dynamics Reinforcement LearningZhenghai Xue, Lang Feng, Jiacheng Xu, Kang Kang 等ICML 2025
它引用的顶会 Paper7
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Discriminator-Weighted Offline Imitation Learning from Suboptimal DemonstrationsHaoran Xu, Xianyuan Zhan, Honglei Yin, Huiling QinICML 2022 · 被引用 105 次
- Uncertainty-Aware Action Advising for Deep Reinforcement Learning AgentsFelipe Leno da Silva, Pablo Hernandez-Leal, Bilal Kartal, Matthew E. TaylorAAAI 2020 · 被引用 84 次
- Efficient Learning of Safe Driving Policy via Human-AI Copilot OptimizationQuanyi Li, Zhenghao Peng, Bolei ZhouICLR 2022 · 被引用 80 次
- Confidence-Aware Imitation Learning from Demonstrations with Varying OptimalitySongyuan Zhang, Zhangjie Cao, Dorsa Sadigh, Yanan SuiNeurIPS 2021 · 被引用 73 次
相关 Paper
- Knowledge Transfer in Multi-Task Deep Reinforcement Learning for Continuous ControlZhiyuan Xu, Kun Wu, Zhengping Che, Jian Tang 等NeurIPS 2020 · 被引用 58 次
- TGRL: An Algorithm for Teacher Guided Reinforcement LearningIdan Shenfeld, Zhang-Wei Hong, Aviv Tamar, Pulkit AgrawalICML 2023 · 被引用 22 次
- Policy Expansion for Bridging Offline-to-Online Reinforcement LearningHaichao Zhang, Wei Xu, Haonan YuICLR 2023 · 被引用 5 次
- Adaptive Policy Learning for Offline-to-Online Reinforcement LearningHan Zheng, Xufang Luo, Pengfei Wei, Xuan Song 等AAAI 2023 · 被引用 47 次
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 被引用 380 次
