Near-optimal Conservative Exploration in Reinforcement Learning under Episode-wise Constraints
Donghao Li, Ruiquan Huang, Cong Shen, Jing Yang
摘要
This paper investigates conservative exploration in reinforcement learning where the performance of the learning agent is guaranteed to be above a certain threshold throughout the learning process. It focuses on the tabular episodic Markov Decision Process (MDP) setting that has finite states and actions. With the knowledge of an existing safe baseline policy, an algorithm termed as StepMix is proposed to balance the exploitation and exploration while ensuring that the conservative constraint is never violated in each episode with high probability. StepMix features a unique design of a mixture policy that adaptively and smoothly interpolates between the baseline policy and the optimistic policy. Theoretical analysis shows that StepMix achieves near-optimal regret order as in the constraint-free setting, indicating that obeying the stringent episode-wise conservative constraint does not compromise the learning performance. Besides, a randomization-based EpsMix algorithm is also proposed and shown to achieve the same performance as StepMix. The algorithm design and theoretical analysis are further extended to the setting where the baseline policy is not given a priori but must be learned from an offline dataset, and it is proved that similar conservative guarantee and regret can be achieved if the offline dataset is sufficiently large. Experiment results corroborate the theoretical analysis and demonstrate the effectiveness of the proposed conservative exploration strategies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper19
- Natural Policy Gradient Primal-Dual Method for Constrained Markov Decision ProcessesDongsheng Ding, Kaiqing Zhang, Tamer Basar, Mihailo R. JovanovicNeurIPS 2020 · 被引用 252 次
- Policy Finetuning: Bridging Sample-Efficient Offline and Online Reinforcement LearningTengyang Xie, Nan Jiang, Huan Wang, Caiming Xiong 等NeurIPS 2021 · 被引用 207 次
- Fast active learning for pure exploration in reinforcement learningPierre Ménard, Omar Darwiche Domingues, Anders Jonsson, Emilie Kaufmann 等ICML 2021 · 被引用 110 次
- Learning Policies with Zero or Bounded Constraint Violation for Constrained MDPsTao Liu, Ruida Zhou, Dileep Kalathil, Panganamala R. Kumar 等NeurIPS 2021 · 被引用 110 次
- Safe Reinforcement Learning via Curriculum InductionMatteo Turchetta, Andrey Kolobov, Shital Shah, Andreas Krause 等NeurIPS 2020 · 被引用 109 次
相关 Paper
- A Near-Optimal Algorithm for Safe Reinforcement Learning Under Instantaneous Hard ConstraintsMing Shi, Yingbin Liang, Ness B. ShroffICML 2023 · 被引用 18 次
- Constraints Penalized Q-learning for Safe Offline Reinforcement LearningHaoran Xu, Xianyuan Zhan, Xiangyu ZhuAAAI 2022 · 被引用 127 次
- Reward-agnostic Fine-tuning: Provable Statistical Benefits of Hybrid Reinforcement LearningGen Li, Wenhao Zhan, Jason D. Lee, Yuejie Chi 等NeurIPS 2023 · 被引用 22 次
- DOPE: Doubly Optimistic and Pessimistic Exploration for Safe Reinforcement LearningArchana Bura, Aria HasanzadeZonuzy, Dileep Kalathil, Srinivas Shakkottai 等NeurIPS 2022 · 被引用 48 次
- RAMBO-RL: Robust Adversarial Model-Based Offline Reinforcement LearningMarc Rigter, Bruno Lacerda, Nick HawesNeurIPS 2022 · 被引用 168 次
