Bilinear Exponential Family of MDPs: Frequentist Regret Bound with Tractable Exploration & Planning
Reda Ouhamma, Debabrota Basu, Odalric Maillard
摘要
We study the problem of episodic reinforcement learning in continuous stateaction spaces with unknown rewards and transitions. Specifically, we consider the setting where the rewards and transitions are modeled using parametric bilinear exponential families. We propose an algorithm, BEF-RLSVI, that a) uses penalized maximum likelihood estimators to learn the unknown parameters, b) injects a calibrated Gaussian noise in the parameter of rewards to ensure exploration, and c) leverages linearity of the exponential family with respect to an underlying RKHS to perform tractable planning. We further provide a frequentist regret analysis of BEF-RLSVI that yields an upper bound of Õ( √ d 3 H 3 K), where d is the dimension of the parameters, H is the episode length, and K is the number of episodes. Our analysis improves the existing bounds for the bilinear exponential family of MDPs by √ H and removes the handcrafted clipping deployed in existing RLSVI-type algorithms. Our regret bound is order-optimal with respect to H and K. * https://redaouhamma.github.io/ Preprint. Under review. Table 1: A comparison of RL Algorithms for MDPs with functional representations. Algorithm Regret Tractable Tractable Free of Model, assumptions exploration planning clipping Thompson sampling √ d 2 H 3 K ✗ ✓ N.A Gaussian P [RZSD21] (Bayesian) Known rewards EXP-UCRL √ d 2 H 4 K ✗ ✗ N.A Bilinear Exp Family (BEF) [CGM21] (Frequentist) known rewards SMRL [LLS + 21] √ d 2 H 4 K ✗ ✗ N.A BEF, known rewards UCRL-VTR [AJS + 20] √ d 2 H 4 K ✗ ✗ N.A Linear mixture model F -PHE-LSVI [ICN + 21] poly(dEH) √ KH ✓ ✗ ✗ Eluder dimension, Tabular PHE-LSVI (linear-RL) √ d 3 H 4 K Anti-concentration UC-MatrixRL [YW20] √ d 2 H 5 K ✗ ✗ N.A Linear factor MDP OPT-RLSVI [ZBB + 20] √ d 4 H 5 K ✓ ✓ ✗ Linear V BEF-RLSVI (this work) √ d 3 H 3 K ✓ ✓ ✓ Bilinear Exp Family
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Provable and Practical: Efficient Exploration in Reinforcement Learning via Langevin Monte CarloHaque Ishfaq, Qingfeng Lan, Pan Xu, A. Rupam Mahmood 等ICLR 2024 · 被引用 33 次
- Diffusion Spectral Representation for Reinforcement LearningDmitry Shribak, Chen-Xiao Gao, Yitong Li, Chenjun Xiao 等NeurIPS 2024 · 被引用 18 次
- Efficient Preference-Based Reinforcement Learning: Randomized Exploration meets Experimental DesignAndreas Schlaginhaufen, Reda Ouhamma, Maryam KamgarpourNeurIPS 2025 · 被引用 4 次
- Performative Policy Gradient: Optimality in Performative Reinforcement LearningDebabrota Basu, Udvas Das, Brahim Driss, Uddalak MukherjeeICML 2026 · 被引用 2 次
- Asymptotically Optimal Sequential Testing with Markovian DataAlhad Sethi, SOFIA SAGAR KAVALI, Shubhada Agrawal, Debabrota Basu 等ICML 2026
它引用的顶会 Paper11
- Model-Based Reinforcement Learning with Value-Targeted RegressionAlex Ayoub, Zeyu Jia, Csaba Szepesvári, Mengdi Wang 等ICML 2020 · 被引用 324 次
- Reinforcement Learning in Feature Space: Matrix Bandit, Kernels, and Regret BoundLin Yang, Mengdi WangICML 2020 · 被引用 308 次
- Bilinear Classes: A Structural Framework for Provable Generalization in RLSimon S. Du, Sham M. Kakade, Jason D. Lee, Shachar Lovett 等ICML 2021 · 被引用 207 次
- Learning with Good Feature Representations in Bandits and in RL with a Generative ModelTor Lattimore, Csaba Szepesvári, Gellért WeiszICML 2020 · 被引用 181 次
- Reinforcement Learning with General Value Function Approximation: Provably Efficient Approach via Bounded Eluder DimensionRuosong Wang, Ruslan Salakhutdinov, Lin F. YangNeurIPS 2020 · 被引用 168 次
相关 Paper
- Randomized Exploration for Reinforcement Learning with Multinomial Logistic Function ApproximationWooseong Cho, Taehyun Hwang, Joongkyu Lee, Min-hwan OhNeurIPS 2024 · 被引用 7 次
- Risk-Sensitive Reinforcement Learning: Near-Optimal Risk-Sample Tradeoff in RegretYingjie Fei, Zhuoran Yang, Yudong Chen, Zhaoran Wang 等NeurIPS 2020 · 被引用 87 次
- Learning Adversarial Linear Mixture Markov Decision Processes with Bandit Feedback and Unknown TransitionCanzhe Zhao, Ruofeng Yang, Baoxiang Wang, Shuai LiICLR 2023
- Exponential Family Model-Based Reinforcement Learning via Score MatchingGene Li, Junbo Li, Anmol Kabra, Nati Srebro 等NeurIPS 2022 · 被引用 5 次
- Model-based Reinforcement Learning for Continuous Control with Posterior SamplingYing Fan, Yifei MingICML 2021 · 被引用 25 次
