High Confidence Generalization for Reinforcement Learning
James E. Kostas, Yash Chandak, Scott M. Jordan, Georgios Theocharous, Philip S. Thomas
Abstract
We present several classes of reinforcement learning algorithms that safely generalize to Markov decision processes (MDPs) not seen during training. Specifically, we study the setting in which some set of MDPs is accessible for training. For various definitions of safety, our algorithms give probabilistic guarantees that agents can safely generalize to MDPs that are sampled from the same distribution but are not necessarily in the training set. These algorithms are a type of Seldonian algorithm (Thomas et al., 2019), which is a class of machine learning algorithms that return models with probabilistic safety guarantees for user-specified definitions of safety.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa92aac2-0259-44f5-bcbf-1812db8c867eCited by top-tier papers1
Ask how each one uses itBuilds on1
Related papers
- Towards Safe Policy Improvement for Non-Stationary MDPsYash Chandak, Scott M. Jordan, Georgios Theocharous, Martha White et al.NeurIPS 2020 · 47 citations
- Security Analysis of Safe and Seldonian Reinforcement Learning AlgorithmsPinar Ozisik, Philip S. ThomasNeurIPS 2020 · 4 citations
- Safe Reinforcement Learning Using Advantage-Based InterventionNolan Wagener, Byron Boots, Ching-An ChengICML 2021 · 66 citations
- Probabilistic Shielding for Safe Reinforcement LearningEdwin Hamel-De le Court, Francesco Belardinelli, Alexander W. GoodallAAAI 2025 · 7 citations
- Safe Reinforcement Learning in Constrained Markov Decision ProcessesAkifumi Wachi, Yanan SuiICML 2020 · 190 citations
