Regularization Guarantees Generalization in Bayesian Reinforcement Learning through Algorithmic Stability
Aviv Tamar, Daniel Soudry, Ev Zisselman
Abstract
In the Bayesian reinforcement learning (RL) setting, a prior distribution over the unknown problem parameters -- the rewards and transitions -- is assumed, and a policy that optimizes the (posterior) expected return is sought. A common approximation, which has been recently popularized as meta-RL, is to train the agent on a sample of N problem instances from the prior, with the hope that for large enough N, good generalization behavior to an unseen test instance will be obtained. In this work, we study generalization in Bayesian RL under the probably approximately correct (PAC) framework, using the method of algorithmic stability. Our main contribution is showing that by adding regularization, the optimal policy becomes uniformly stable in an appropriate sense. Most stability results in the literature build on strong convexity of the regularized loss -- an approach that is not suitable for RL as Markov decision processes (MDPs) are not convex. Instead, building on recent results of fast convergence rates for mirror descent in regularized MDPs, we show that regularized MDPs satisfy a certain quadratic growth criterion, which is sufficient to establish stability. This result, which may be of independent interest, allows us to study the effect of regularization on generalization in the Bayesian RL setting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d838ac73-9364-401a-87f0-e92a43383afdCited by top-tier papers4
- Explore to Generalize in Zero-Shot RLEv Zisselman, Itai Lavie, Daniel Soudry, Aviv TamarNeurIPS 2023 · 26 citations
- Meta Reinforcement Learning with Finite Training Tasks - a Density Estimation ApproachZohar Rimon, Aviv Tamar, Gilad AdlerNeurIPS 2022 · 9 citations
- Test-Time Regret Minimization in Meta Reinforcement LearningMirco Mutti, Aviv TamarICML 2024 · 4 citations
- Stability beyond Bounded Differences: Sharp Generalization Bounds under Finite MomentsQianqian Lei, Soham Bonnerjee, Yuefeng Han, Wei Biao WuICML 2026 · 1 citation
Builds on7
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 911 citations
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 685 citations
- Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPsLior Shani, Yonathan Efroni, Shie MannorAAAI 2020 · 201 citations
- PACOH: Bayes-Optimal Meta-Learning with PAC-GuaranteesJonas Rothfuss, Vincent Fortuin, Martin Josifoski, Andreas KrauseICML 2021 · 136 citations
- Meta-Learning with Fewer Tasks through Task InterpolationHuaxiu Yao, Linjun Zhang, Chelsea FinnICLR 2022 · 66 citations
Related papers
- Generalization Bounds for Meta-Learning via PAC-Bayes and Uniform StabilityAlec Farid, Anirudha MajumdarNeurIPS 2021 · 46 citations
- On the Stability and Generalization of Meta-LearningYunjuan Wang, Raman AroraNeurIPS 2024 · 12 citations
- Multi-Agent Meta-Reinforcement Learning: Sharper Convergence Rates with Task SimilarityWeichao Mao, Haoran Qiu, Chen Wang, Hubertus Franke et al.NeurIPS 2023 · 17 citations
- Bayesian decision-making under misspecified priors with applications to meta-learningMax Simchowitz, Christopher Tosh, Akshay Krishnamurthy, Daniel J. Hsu et al.NeurIPS 2021 · 57 citations
- Meta-trained agents implement Bayes-optimal agentsVladimir Mikulik, Grégoire Delétang, Tom McGrath, Tim Genewein et al.NeurIPS 2020 · 56 citations
