STEERING : Stein Information Directed Exploration for Model-Based Reinforcement Learning
Souradip Chakraborty, Amrit S. Bedi, Alec Koppel, Mengdi Wang, Furong Huang, Dinesh Manocha
Abstract
Directed Exploration is a crucial challenge in reinforcement learning (RL), especially when rewards are sparse. Information-directed sampling (IDS), which optimizes the information ratio, seeks to do so by augmenting regret with information gain. However, estimating information gain is computationally intractable or relies on restrictive assumptions which prohibit its use in many practical instances. In this work, we posit an alternative exploration incentive in terms of the integral probability metric (IPM) between a current estimate of the transition model and the unknown optimal, which under suitable conditions, can be computed in closed form with the kernelized Stein discrepancy (KSD). Based on KSD, we develop a novel algorithm : STEin information dirEcted exploration for model-based Reinforcement LearnING. To enable its derivation, we develop fundamentally new variants of KSD for discrete conditional distributions. We further establish that archives sublinear Bayesian regret, improving upon prior learning rates of information-augmented MBRL. Experimentally, we show that the proposed algorithm is computationally affordable and outperforms several prior approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 00c41aee-1fd3-4311-8ead-a8fa85ff5051Cited by top-tier papers4
- Optimistic Active Exploration of Dynamical SystemsBhavya Sukhija, Lenart Treven, Cansu Sancaktar, Sebastian Blaes et al.NeurIPS 2023 · 42 citations
- PARL: A Unified Framework for Policy Alignment in Reinforcement Learning from Human FeedbackSouradip Chakraborty, Amrit Singh Bedi, Alec Koppel, Huazheng Wang et al.ICLR 2024 · 42 citations
- Information-Directed Pessimism for Offline Reinforcement LearningAlec Koppel, Sujay Bhatt, Jiacheng Guo, Joe Eappen et al.ICML 2024 · 3 citations
- Maximum Total Correlation Reinforcement LearningBang You, Puze Liu, Huaping Liu, Jan Peters et al.ICML 2025
Builds on8
- Model-Based Reinforcement Learning with Value-Targeted RegressionAlex Ayoub, Zeyu Jia, Csaba Szepesvári, Mengdi Wang et al.ICML 2020 · 324 citations
- Reinforcement Learning in Feature Space: Matrix Bandit, Kernels, and Regret BoundLin Yang, Mengdi WangICML 2020 · 308 citations
- Maximum Entropy RL (Provably) Solves Some Robust RL ProblemsBenjamin Eysenbach, Sergey LevineICLR 2022 · 244 citations
- Reinforcement Learning with Sparse Rewards using Guidance from Offline DemonstrationDesik Rengarajan, Gargi Vaidya, Akshay Sarvesh, Dileep M. Kalathil et al.ICLR 2022 · 86 citations
- Regret Bounds for Information-Directed Reinforcement LearningBotao Hao, Tor LattimoreNeurIPS 2022 · 31 citations
Related papers
- Posterior Coreset Construction with Kernelized Stein Discrepancy for Model-Based Reinforcement LearningSouradip Chakraborty, Amrit Singh Bedi, Pratap Tokekar, Alec Koppel et al.AAAI 2023 · 10 citations
- IMED-RL: Regret optimal learning of ergodic Markov decision processesFabien Pesquerel, Odalric-Ambrym MaillardNeurIPS 2022 · 12 citations
- Implicit Generative Modeling for Efficient ExplorationNeale Ratzlaff, Qinxun Bai, Fuxin Li, Wei XuICML 2020 · 15 citations
- Kernel-Based Reinforcement Learning: A Finite-Time AnalysisOmar Darwiche Domingues, Pierre Ménard, Matteo Pirotta, Emilie Kaufmann et al.ICML 2021 · 24 citations
- Information Directed Reward Learning for Reinforcement LearningDavid Lindner, Matteo Turchetta, Sebastian Tschiatschek, Kamil Ciosek et al.NeurIPS 2021 · 27 citations
