Constrained Policy Optimization via Bayesian World Models
Yarden As, Ilnura Usmanova, Sebastian Curi, Andreas Krause
摘要
Improving sample-efficiency and safety are crucial challenges when deploying reinforcement learning in high-stakes real world applications. We propose LAMBDA, a novel model-based approach for policy optimization in safety critical tasks modeled via constrained Markov decision processes. Our approach utilizes Bayesian world models, and harnesses the resulting uncertainty to maximize optimistic upper bounds on the task objective, as well as pessimistic upper bounds on the safety constraints. We demonstrate LAMBDA's state of the art performance on the Safety-Gym benchmark suite in terms of sample efficiency and constraint violation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- Constrained Variational Policy Optimization for Safe Reinforcement LearningZuxin Liu, Zhepeng Cen, Vladislav Isenbaev, Wei Liu 等ICML 2022 · 被引用 112 次
- Enforcing Hard Constraints with Soft Barriers: Safe Reinforcement Learning in Unknown Stochastic EnvironmentsYixuan Wang, Simon Sinong Zhan, Ruochen Jiao, Zhilu Wang 等ICML 2023 · 被引用 81 次
- Towards Safe Reinforcement Learning with a Safety Editor PolicyHaonan Yu, Wei Xu, Haichao ZhangNeurIPS 2022 · 被引用 50 次
- SafeDreamer: Safe Reinforcement Learning with World ModelsWeidong Huang, Jiaming Ji, Chunhe Xia, Borong Zhang 等ICLR 2024 · 被引用 46 次
- Safe Offline Reinforcement Learning with Real-Time Budget ConstraintsQian Lin, Bo Tang, Zifan Wu, Chao Yu 等ICML 2023 · 被引用 31 次
它引用的顶会 Paper6
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Responsive Safety in Reinforcement Learning by PID Lagrangian MethodsAdam Stooke, Joshua Achiam, Pieter AbbeelICML 2020 · 被引用 403 次
- Efficient Model-Based Reinforcement Learning through Optimistic Policy Search and PlanningSebastian Curi, Felix Berkenkamp, Andreas KrauseNeurIPS 2020 · 被引用 120 次
- Safe Reinforcement Learning via Curriculum InductionMatteo Turchetta, Andrey Kolobov, Shital Shah, Andreas Krause 等NeurIPS 2020 · 被引用 109 次
- Conservative Safety Critics for ExplorationHomanga Bharadhwaj, Aviral Kumar, Nicholas Rhinehart, Sergey Levine 等ICLR 2021 · 被引用 32 次
相关 Paper
- Model-based Safe Deep Reinforcement Learning via a Constrained Proximal Policy Optimization AlgorithmAshish Kumar Jayant, Shalabh BhatnagarNeurIPS 2022 · 被引用 84 次
- Enhancing Efficiency of Safe Reinforcement Learning via Sample ManipulationShangding Gu, Laixi Shi, Yuhao Ding, Alois Knoll 等NeurIPS 2024 · 被引用 14 次
- Conservative and Adaptive Penalty for Model-Based Safe Reinforcement LearningYecheng Jason Ma, Andrew Shen, Osbert Bastani, Dinesh JayaramanAAAI 2022 · 被引用 32 次
- Learning with Safety Constraints: Sample Complexity of Reinforcement Learning for Constrained MDPsAria HasanzadeZonuzy, Archana Bura, Dileep M. Kalathil, Srinivas ShakkottaiAAAI 2021 · 被引用 46 次
- ActSafe: Active Exploration with Safety Constraints for Reinforcement LearningYarden As, Bhavya Sukhija, Lenart Treven, Carmelo Sferrazza 等ICLR 2025
