Exponential Family Model-Based Reinforcement Learning via Score Matching
Gene Li, Junbo Li, Anmol Kabra, Nati Srebro, Zhaoran Wang, Zhuoran Yang
Abstract
We propose an optimistic model-based algorithm, dubbed SMRL, for finite-horizon episodic reinforcement learning (RL) when the transition model is specified by exponential family distributions with parameters and the reward is bounded and known. SMRL uses score matching, an unnormalized density estimation technique that enables efficient estimation of the model parameter by ridge regression. Under standard regularity assumptions, SMRL achieves online regret, where is the length of each episode and is the total number of interactions (ignoring polynomial dependence on structural scale parameters).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f40ad820-32b2-47cc-9071-4e726a98008fCited by top-tier papers3
- Learning a Diffusion Model Policy from Rewards via Q-Score MatchingMichael Psenka, Alejandro Escontrela, Pieter Abbeel, Yi MaICML 2024 · 90 citations
- Bilinear Exponential Family of MDPs: Frequentist Regret Bound with Tractable Exploration & PlanningReda Ouhamma, Debabrota Basu, Odalric MaillardAAAI 2023 · 14 citations
- Provably Efficient Reinforcement Learning with Multinomial Logit Function ApproximationLong-Fei Li, Yu-Jie Zhang, Peng Zhao, Zhi-Hua ZhouNeurIPS 2024 · 11 citations
Builds on7
- Model-Based Reinforcement Learning with Value-Targeted RegressionAlex Ayoub, Zeyu Jia, Csaba Szepesvári, Mengdi Wang et al.ICML 2020 · 324 citations
- Provably Efficient Exploration in Policy OptimizationQi Cai, Zhuoran Yang, Chi Jin, Zhaoran WangICML 2020 · 304 citations
- Learning Near Optimal Policies with Low Inherent Bellman ErrorAndrea Zanette, Alessandro Lazaric, Mykel J. Kochenderfer, Emma BrunskillICML 2020 · 238 citations
- Naive Exploration is Optimal for Online LQRMax Simchowitz, Dylan J. FosterICML 2020 · 209 citations
- Learning with Good Feature Representations in Bandits and in RL with a Generative ModelTor Lattimore, Csaba Szepesvári, Gellért WeiszICML 2020 · 181 citations
Related papers
- Randomized Exploration for Reinforcement Learning with Multinomial Logistic Function ApproximationWooseong Cho, Taehyun Hwang, Joongkyu Lee, Min-hwan OhNeurIPS 2024 · 7 citations
- Kernel-Based Reinforcement Learning: A Finite-Time AnalysisOmar Darwiche Domingues, Pierre Ménard, Matteo Pirotta, Emilie Kaufmann et al.ICML 2021 · 24 citations
- Optimistic Posterior Sampling for Reinforcement Learning with Few Samples and Tight GuaranteesDaniil Tiapkin, Denis Belomestny, Daniele Calandriello, Eric Moulines et al.NeurIPS 2022 · 16 citations
- Optimistic Policy Optimization with Bandit FeedbackLior Shani, Yonathan Efroni, Aviv Rosenberg, Shie MannorICML 2020 · 100 citations
- Online Reinforcement Learning with Uncertain Episode LengthsDebmalya Mandal, Goran Radanovic, Jiarui Gan, Adish Singla et al.AAAI 2023 · 9 citations
