Posterior Coreset Construction with Kernelized Stein Discrepancy for Model-Based Reinforcement Learning
Souradip Chakraborty, Amrit Singh Bedi, Pratap Tokekar, Alec Koppel, Brian M. Sadler, Furong Huang, Dinesh Manocha
Abstract
Model-based approaches to reinforcement learning (MBRL) exhibit favorable performance in practice, but their theoretical guarantees in large spaces are mostly restricted to the setting when transition model is Gaussian or Lipschitz, and demands a posterior estimate whose representational complexity grows unbounded with time. In this work, we develop a novel MBRL method (i) which relaxes the assumptions on the target transition model to belong to a generic family of mixture models; (ii) is applicable to large-scale training by incorporating a compression step such that the posterior estimate consists of a Bayesian coreset of only statistically significant past state-action pairs; and (iii) exhibits a sublinear Bayesian regret. To achieve these results, we adopt an approach based upon Stein's method, which, under a smoothness condition on the constructed posterior and target, allows distributional distance to be evaluated in closed form as the kernelized Stein discrepancy (KSD). The aforementioned compression step is then computed in terms of greedily retaining only those samples which are more than a certain KSD away from the previous model estimate. Experimentally, we observe that this approach is competitive with several state-of-the-art RL methodologies, and can achieve up-to 50 percent reduction in wall clock time in some continuous control environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b9c83b4-975e-4339-bf64-d3594cbde754Cited by top-tier papers4
- PARL: A Unified Framework for Policy Alignment in Reinforcement Learning from Human FeedbackSouradip Chakraborty, Amrit Singh Bedi, Alec Koppel, Huazheng Wang et al.ICLR 2024 · 42 citations
- STEERING : Stein Information Directed Exploration for Model-Based Reinforcement LearningSouradip Chakraborty, Amrit S. Bedi, Alec Koppel, Mengdi Wang et al.ICML 2023 · 9 citations
- Improved Bayesian Regret Bounds for Thompson Sampling in Reinforcement LearningAhmadreza Moradipari, Mohammad Pedramfar, Modjtaba Shokrian Zini, Vaneet AggarwalNeurIPS 2023 · 8 citations
- Information-Directed Pessimism for Offline Reinforcement LearningAlec Koppel, Sujay Bhatt, Jiacheng Guo, Joe Eappen et al.ICML 2024 · 3 citations
Builds on6
- Model-Based Reinforcement Learning with Value-Targeted RegressionAlex Ayoub, Zeyu Jia, Csaba Szepesvári, Mengdi Wang et al.ICML 2020 · 324 citations
- Reinforcement Learning in Feature Space: Matrix Bandit, Kernels, and Regret BoundLin Yang, Mengdi WangICML 2020 · 308 citations
- Learning Near Optimal Policies with Low Inherent Bellman ErrorAndrea Zanette, Alessandro Lazaric, Mykel J. Kochenderfer, Emma BrunskillICML 2020 · 238 citations
- Variational Bayesian Reinforcement Learning with Regret BoundsBrendan O'DonoghueNeurIPS 2021 · 48 citations
- Model-based Reinforcement Learning for Continuous Control with Posterior SamplingYing Fan, Yifei MingICML 2021 · 25 citations
Related papers
- Model-based micro-data reinforcement learning: what are the crucial model properties and which model to choose?Balázs Kégl, Gabriel Hurtado, Albert ThomasICLR 2021 · 13 citations
- Debiased Distribution CompressionLingxiao Li, Raaz Dwivedi, Lester MackeyICML 2024 · 7 citations
- Minimum Description Length ControlTed Moskovitz, Ta-Chu Kao, Maneesh Sahani, Matt M. BotvinickICLR 2023 · 76 citations
- Near-Optimal Reward-Free Exploration for Linear Mixture MDPs with Plug-in SolverXiaoyu Chen, Jiachen Hu, Lin Yang, Liwei WangICLR 2022 · 14 citations
- Regularized Offline Policy Optimization with Posterior Hybrid Bayesian BeliefHongqiang Lin, Pengfei Wang, Nenggan ZhengICML 2026 · 1 citation
