Blending Imitation and Reinforcement Learning for Robust Policy Improvement
Xuefeng Liu, Takuma Yoneda, Rick Stevens, Matthew R. Walter, Yuxin Chen
摘要
While reinforcement learning (RL) has shown promising performance, its sample complexity continues to be a substantial hurdle, restricting its broader application across a variety of domains. Imitation learning (IL) utilizes oracles to improve sample efficiency, yet it is often constrained by the quality of the oracles deployed. To address the demand for robust policy improvement in real-world scenarios, we introduce a novel algorithm, Robust Policy Improvement (RPI), which actively interleaves between IL and RL based on an online estimate of their performance. RPI draws on the strengths of IL, using oracle queries to facilitate exploration-an aspect that is notably challenging in sparse-reward RL-particularly during the early stages of learning. As learning unfolds, RPI gradually transitions to RL, effectively treating the learned policy as an improved oracle. This algorithm is capable of learning from and improving upon a diverse set of black-box oracles. Integral to RPI are Robust Active Policy Selection (RAPS) and Robust Policy Gradient (RPG), both of which reason over whether to perform state-wise imitation from the oracles or learn from its own value function when the learner's performance surpasses that of the oracles in a specific state. Empirical evaluations and theoretical analysis validate that RPI excels in comparison to existing state-ofthe-art methods, showing superior performance across various domains. Please checkout our website 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Entropy-Reinforced Planning with Large Language Models for Drug DiscoveryXuefeng Liu, Chih-chan Tien, Peng Ding, Songhao Jiang 等ICML 2024 · 被引用 7 次
- Oracle-Efficient Reinforcement Learning for Max Value EnsemblesMarcel Hussing, Michael Kearns, Aaron Roth, Sikata Bela Sengupta 等NeurIPS 2024 · 被引用 2 次
- DexTrack: Towards Generalizable Neural Tracking Control for Dexterous Manipulation from Human ReferencesXueyi Liu, Jianibieke Adalibieke, Qianwei Han, Yuzhe Qin 等ICLR 2025
- Reinforcement Learning from Imperfect Corrective Actions and Proxy RewardsZhaohui Jiang, Xuening Feng, Paul Weng, Yifei Zhu 等ICLR 2025
它引用的顶会 Paper10
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- AlphaHoldem: High-Performance Artificial Intelligence for Heads-Up No-Limit Poker via End-to-End Reinforcement LearningEnmin Zhao, Renye Yan, Jinqiu Li, Kai Li 等AAAI 2022 · 被引用 63 次
- Imitation Learning by Estimating Expertise of DemonstratorsMark Beliaev, Andy Shih, Stefano Ermon, Dorsa Sadigh 等ICML 2022 · 被引用 60 次
- Policy Improvement via Imitation of Multiple OraclesChing-An Cheng, Andrey Kolobov, Alekh AgarwalNeurIPS 2020 · 被引用 36 次
- Active Offline Policy SelectionKsenia Konyushkova, Yutian Chen, Thomas Paine, Çaglar Gülçehre 等NeurIPS 2021 · 被引用 35 次
相关 Paper
- Active Policy Improvement from Multiple Black-box OraclesXuefeng Liu, Takuma Yoneda, Chaoqi Wang, Matthew R. Walter 等ICML 2023 · 被引用 13 次
- Local policy search with Bayesian optimizationSarah Müller, Alexander von Rohr, Sebastian TrimpeNeurIPS 2021 · 被引用 67 次
- Breaking the Computational Barrier: Provably Efficient Actor–Critic for Low-Rank MDPsRuiquan Huang, Donghao Li, Yingbin LIANG, Jing YangICML 2026
- A Few Expert Queries Suffices for Sample-Efficient RL with Resets and Linear Value ApproximationPhilip Amortila, Nan Jiang, Dhruv Madeka, Dean P. FosterNeurIPS 2022 · 被引用 6 次
- RLIF: Interactive Imitation Learning as Reinforcement LearningJianlan Luo, Perry Dong, Yuexiang Zhai, Yi Ma 等ICLR 2024 · 被引用 31 次
