Multi-Objective SPIBB: Seldonian Offline Policy Improvement with Safety Constraints in Finite MDPs
Harsh Satija, Philip S. Thomas, Joelle Pineau, Romain Laroche
Abstract
We study the problem of Safe Policy Improvement (SPI) under constraints in the offline Reinforcement Learning (RL) setting. We consider the scenario where: (i) we have a dataset collected under a known baseline policy, (ii) multiple reward signals are received from the environment inducing as many objectives to optimize. We present an SPI formulation for this RL setting that takes into account the preferences of the algorithm's user for handling the trade-offs for different reward signals while ensuring that the new policy performs at least as well as the baseline policy along each individual objective. We build on traditional SPI algorithms and propose a novel method based on Safe Policy Iteration with Baseline Bootstrapping (SPIBB, Laroche et al., 2019) that provides high probability guarantees on the performance of the agent in the true environment. We show the effectiveness of our method on a synthetic grid-world safety task as well as in a real-world critical care context to learn a policy for the administration of IV fluids and vasopressors to treat sepsis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d477e39d-d622-4470-8b18-1f51160f93e1Cited by top-tier papers7
- Direction-oriented Multi-objective Learning: Simple and Provable Stochastic AlgorithmsPeiyao Xiao, Hao Ban, Kaiyi JiNeurIPS 2023 · 46 citations
- Adaptive Stochastic Gradient Algorithm for Black-box Multi-Objective LearningFeiyang Ye, Yueming Lyu, Xuehao Wang, Yu Zhang et al.ICLR 2024 · 5 citations
- Behavior Prior Representation learning for Offline Reinforcement LearningHongyu Zang, Xin Li, Jie Yu, Chen Liu et al.ICLR 2023 · 3 citations
- A Provable Approach for End-to-End Safe Reinforcement LearningAkifumi Wachi, Kohei Miyaguchi, Takumi Tanabe, Rei Sato et al.NeurIPS 2025 · 2 citations
- Scaling Pareto-Efficient Decision Making via Offline Multi-Objective RLBaiting Zhu, Meihua Dang, Aditya GroverICLR 2023 · 1 citation
Builds on2
- A distributional view on multi-objective policy optimizationAbbas Abdolmaleki, Sandy H. Huang, Leonard Hasenclever, Michael Neunert et al.ICML 2020 · 93 citations
- Clinician-in-the-Loop Decision Making: Reinforcement Learning with Near-Optimal Set-Valued PoliciesShengpu Tang, Aditya Modi, Michael W. Sjoding, Jenna WiensICML 2020 · 35 citations
Related papers
- Scalable Safe Policy Improvement via Monte Carlo Tree SearchAlberto Castellini, Federico Bianchi, Edoardo Zorzi, Thiago D. Simão et al.ICML 2023 · 9 citations
- Offline Guarded Safe Reinforcement Learning for Medical Treatment Optimization StrategiesRunze Yan, Xun Shen, Akifumi Wachi, Sebastien Gros et al.NeurIPS 2025 · 7 citations
- Constraints Penalized Q-learning for Safe Offline Reinforcement LearningHaoran Xu, Xianyuan Zhan, Xiangyu ZhuAAAI 2022 · 127 citations
- Accelerating Safe Reinforcement Learning with Constraint-mismatched Baseline PoliciesTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICML 2021 · 20 citations
- COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction EstimationJongmin Lee, Cosmin Paduraru, Daniel J. Mankowitz, Nicolas Heess et al.ICLR 2022 · 84 citations
