TGRL: An Algorithm for Teacher Guided Reinforcement Learning
Idan Shenfeld, Zhang-Wei Hong, Aviv Tamar, Pulkit Agrawal
摘要
Learning from rewards (i.e., reinforcement learning or RL) and learning to imitate a teacher (i.e., teacher-student learning) are two established approaches for solving sequential decision-making problems. To combine the benefits of these different forms of learning, it is common to train a policy to maximize a combination of reinforcement and teacher-student learning objectives. However, without a principled method to balance these objectives, prior work used heuristics and problem-specific hyperparameter searches to balance the two objectives. We present a approach, along with an approximate implementation for and balancing when to follow the teacher and when to use rewards. The main idea is to adjust the importance of teacher supervision by comparing the agent's performance to the counterfactual scenario of the agent learning without teacher supervision and only from rewards. If using teacher supervision improves performance, the importance of teacher supervision is increased and otherwise it is decreased. Our method, (TGRL), outperforms strong baselines across diverse domains without hyper-parameter tuning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Provable Partially Observable Reinforcement Learning with Privileged InformationYang Cai, Xiangyu Liu, Argyris Oikonomou, Kaiqing ZhangNeurIPS 2024 · 被引用 22 次
- Privileged Sensing Scaffolds Reinforcement LearningEdward S. Hu, James Springer, Oleh Rybkin, Dinesh JayaramanICLR 2024 · 被引用 21 次
- Blending Imitation and Reinforcement Learning for Robust Policy ImprovementXuefeng Liu, Takuma Yoneda, Rick Stevens, Matthew R. Walter 等ICLR 2024 · 被引用 19 次
- Iterative Regularized Policy Optimization with Imperfect DemonstrationsXudong Gong, Dawei Feng, Kele Xu, Yuanzhao Zhai 等ICML 2024 · 被引用 5 次
- Multi-Agent Guided Policy OptimizationYueheng Li, Guangming Xie, Zongqing LuICLR 2026 · 被引用 4 次
它引用的顶会 Paper7
- Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPsTianwei Ni, Benjamin Eysenbach, Ruslan SalakhutdinovICML 2022 · 被引用 162 次
- Reincarnating Reinforcement Learning: Reusing Prior Computation to Accelerate ProgressRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2022 · 被引用 95 次
- Sequential Causal Imitation Learning with Unobserved ConfoundersDaniel Kumor, Junzhe Zhang, Elias BareinboimNeurIPS 2021 · 被引用 53 次
- Redeeming intrinsic rewards via constrained optimizationEric Chen, Zhang-Wei Hong, Joni Pajarinen, Pulkit AgrawalNeurIPS 2022 · 被引用 48 次
- Sequence Model Imitation Learning with Unobserved ContextsGokul Swamy, Sanjiban Choudhury, J. Andrew Bagnell, Zhiwei Steven WuNeurIPS 2022 · 被引用 39 次
相关 Paper
- Self-Adaptive Imitation Learning: Learning Tasks with Delayed Rewards from Sub-optimal DemonstrationsZhuangdi Zhu, Kaixiang Lin, Bo Dai, Jiayu ZhouAAAI 2022 · 被引用 14 次
- Guarded Policy Optimization with Imperfect Online DemonstrationsZhenghai Xue, Zhenghao Peng, Quanyi Li, Zhihan Liu 等ICLR 2023 · 被引用 4 次
- Reinforcement Learning Guided Semi-Supervised LearningMarzi Heidari, Hanping Zhang, Yuhong GuoNeurIPS 2024 · 被引用 6 次
- Guided Policy Optimization under Partial ObservabilityYueheng Li, Guangming Xie, Zongqing LuICLR 2026 · 被引用 4 次
- Accelerating Safe Reinforcement Learning with Constraint-mismatched Baseline PoliciesTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICML 2021 · 被引用 20 次
