Regularization Guarantees Generalization in Bayesian Reinforcement Learning through Algorithmic Stability
Aviv Tamar, Daniel Soudry, Ev Zisselman
摘要
In the Bayesian reinforcement learning (RL) setting, a prior distribution over the unknown problem parameters -- the rewards and transitions -- is assumed, and a policy that optimizes the (posterior) expected return is sought. A common approximation, which has been recently popularized as meta-RL, is to train the agent on a sample of N problem instances from the prior, with the hope that for large enough N, good generalization behavior to an unseen test instance will be obtained. In this work, we study generalization in Bayesian RL under the probably approximately correct (PAC) framework, using the method of algorithmic stability. Our main contribution is showing that by adding regularization, the optimal policy becomes uniformly stable in an appropriate sense. Most stability results in the literature build on strong convexity of the regularized loss -- an approach that is not suitable for RL as Markov decision processes (MDPs) are not convex. Instead, building on recent results of fast convergence rates for mirror descent in regularized MDPs, we show that regularized MDPs satisfy a certain quadratic growth criterion, which is sufficient to establish stability. This result, which may be of independent interest, allows us to study the effect of regularization on generalization in the Bayesian RL setting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Explore to Generalize in Zero-Shot RLEv Zisselman, Itai Lavie, Daniel Soudry, Aviv TamarNeurIPS 2023 · 被引用 26 次
- Meta Reinforcement Learning with Finite Training Tasks - a Density Estimation ApproachZohar Rimon, Aviv Tamar, Gilad AdlerNeurIPS 2022 · 被引用 9 次
- Test-Time Regret Minimization in Meta Reinforcement LearningMirco Mutti, Aviv TamarICML 2024 · 被引用 4 次
- Stability beyond Bounded Differences: Sharp Generalization Bounds under Finite MomentsQianqian Lei, Soham Bonnerjee, Yuefeng Han, Wei Biao WuICML 2026 · 被引用 1 次
它引用的顶会 Paper7
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 被引用 911 次
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 被引用 685 次
- Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPsLior Shani, Yonathan Efroni, Shie MannorAAAI 2020 · 被引用 201 次
- PACOH: Bayes-Optimal Meta-Learning with PAC-GuaranteesJonas Rothfuss, Vincent Fortuin, Martin Josifoski, Andreas KrauseICML 2021 · 被引用 136 次
- Meta-Learning with Fewer Tasks through Task InterpolationHuaxiu Yao, Linjun Zhang, Chelsea FinnICLR 2022 · 被引用 66 次
相关 Paper
- Generalization Bounds for Meta-Learning via PAC-Bayes and Uniform StabilityAlec Farid, Anirudha MajumdarNeurIPS 2021 · 被引用 46 次
- On the Stability and Generalization of Meta-LearningYunjuan Wang, Raman AroraNeurIPS 2024 · 被引用 12 次
- Multi-Agent Meta-Reinforcement Learning: Sharper Convergence Rates with Task SimilarityWeichao Mao, Haoran Qiu, Chen Wang, Hubertus Franke 等NeurIPS 2023 · 被引用 17 次
- Bayesian decision-making under misspecified priors with applications to meta-learningMax Simchowitz, Christopher Tosh, Akshay Krishnamurthy, Daniel J. Hsu 等NeurIPS 2021 · 被引用 57 次
- Meta-trained agents implement Bayes-optimal agentsVladimir Mikulik, Grégoire Delétang, Tom McGrath, Tim Genewein 等NeurIPS 2020 · 被引用 56 次
