Oracle-Efficient Reinforcement Learning for Max Value Ensembles
Marcel Hussing, Michael Kearns, Aaron Roth, Sikata Bela Sengupta, Jessica Sorrell
摘要
Reinforcement learning (RL) in large or infinite state spaces is notoriously challenging, both theoretically (where worst-case sample and computational complexities must scale with state space cardinality) and experimentally (where function approximation and policy gradient techniques often scale poorly and suffer from instability and high variance). One line of research attempting to address these difficulties makes the natural assumption that we are given a collection of heuristic base or policies upon which we would like to improve in a scalable manner. In this work we aim to compete with the , which at each state follows the action of whichever constituent policy has the highest value. The max-following policy is always at least as good as the best constituent policy, and may be considerably better. Our main result is an efficient algorithm that learns to compete with the max-following policy, given only access to the constituent policies (but not their value functions). In contrast to prior work in similar settings, our theoretical results require only the minimal assumption of an ERM oracle for value function approximation for the constituent policies (and not the global optimal policy or the max-following policy itself) on samplable distributions. We illustrate our algorithm's experimental effectiveness and behavior on several robotic simulation testbeds.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement LearningKimin Lee, Michael Laskin, Aravind Srinivas, Pieter AbbeelICML 2021 · 被引用 239 次
- Dropout Q-Functions for Doubly Efficient Reinforcement LearningTakuya Hiraoka, Takahisa Imagawa, Taisei Hashimoto, Takashi Onishi 等ICLR 2022 · 被引用 157 次
- Discount Factor as a Regularizer in Reinforcement LearningRon Amit, Ron Meir, Kamil CiosekICML 2020 · 被引用 85 次
- Ensemble Bootstrapping for Q-LearningOren Peer, Chen Tessler, Nadav Merlis, Ron MeirICML 2021 · 被引用 56 次
相关 Paper
- Policy Improvement via Imitation of Multiple OraclesChing-An Cheng, Andrey Kolobov, Alekh AgarwalNeurIPS 2020 · 被引用 36 次
- Actor-Critic based Improper Reinforcement LearningMohammadi Zaki, Avi Mohan, Aditya Gopalan, Shie MannorICML 2022 · 被引用 4 次
- Discovering a set of policies for the worst case rewardTom Zahavy, André Barreto, Daniel J. Mankowitz, Shaobo Hou 等ICLR 2021 · 被引用 26 次
- Accelerating Safe Reinforcement Learning with Constraint-mismatched Baseline PoliciesTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICML 2021 · 被引用 20 次
- A Max-Min Entropy Framework for Reinforcement LearningSeungyul Han, Youngchul SungNeurIPS 2021 · 被引用 44 次
