Randomized Exploration for Reinforcement Learning with Multinomial Logistic Function Approximation
Wooseong Cho, Taehyun Hwang, Joongkyu Lee, Min-hwan Oh
摘要
We study reinforcement learning with multinomial logistic (MNL) function approximation where the underlying transition probability kernel of the Markov decision processes (MDPs) is parametrized by an unknown transition core with features of state and action. For the finite horizon episodic setting with inhomogeneous state transitions, we propose provably efficient algorithms with randomized exploration having frequentist regret guarantees. For our first algorithm, , we adapt optimistic sampling to ensure the optimism of the estimated value function with sufficient frequency. We establish that achieves a frequentist regret bound with constant-time computational cost per episode. Here, is the dimension of the transition core, is the horizon length, is the total number of steps, and is a problem-dependent constant. Despite the simplicity and practicality of , its regret bound scales with , which is potentially large in the worst case. To improve the dependence on , we propose , which estimates the value function using the local gradient information of the MNL transition model. We show that its frequentist regret bound is . To the best of our knowledge, these are the first randomized RL algorithms for the MNL transition model that achieve statistical guarantees with constant-time computational cost per episode. Numerical experiments demonstrate the superior performance of the proposed algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple OptionsJoongkyu Lee, Seouh-won Yi, Min-hwan OhNeurIPS 2025 · 被引用 3 次
- Tractable Multinomial Logit Contextual Bandits with Non-Linear UtilitiesTaehyun Hwang, Dahngoon Kim, Min-hwan OhNeurIPS 2025
- Improved Online Confidence Bounds for Multinomial Logistic BanditsJoongkyu Lee, Min-hwan OhICML 2025
- Diversified Multinomial Logit Contextual BanditsHeesang Ann, Taehyun Hwang, Min-hwan OhICLR 2026
它引用的顶会 Paper30
- Model-Based Reinforcement Learning with Value-Targeted RegressionAlex Ayoub, Zeyu Jia, Csaba Szepesvári, Mengdi Wang 等ICML 2020 · 被引用 324 次
- Reinforcement Learning in Feature Space: Matrix Bandit, Kernels, and Regret BoundLin Yang, Mengdi WangICML 2020 · 被引用 308 次
- Provably Efficient Exploration in Policy OptimizationQi Cai, Zhuoran Yang, Chi Jin, Zhaoran WangICML 2020 · 被引用 304 次
- Bellman Eluder Dimension: New Rich Classes of RL Problems, and Sample-Efficient AlgorithmsChi Jin, Qinghua Liu, Sobhan MiryoosefiNeurIPS 2021 · 被引用 264 次
- Is a Good Representation Sufficient for Sample Efficient Reinforcement Learning?Simon S. Du, Sham M. Kakade, Ruosong Wang, Lin F. YangICLR 2020 · 被引用 213 次
相关 Paper
- Model-Based Reinforcement Learning with Multinomial Logistic Function ApproximationTaehyun Hwang, Min-hwan OhAAAI 2023 · 被引用 13 次
- Provably Efficient Reinforcement Learning with Multinomial Logit Function ApproximationLong-Fei Li, Yu-Jie Zhang, Peng Zhao, Zhi-Hua ZhouNeurIPS 2024 · 被引用 11 次
- Nearly Minimax Optimal Reinforcement Learning for Linear Markov Decision ProcessesJiafan He, Heyang Zhao, Dongruo Zhou, Quanquan GuICML 2023 · 被引用 68 次
- Near-Optimal Randomized Exploration for Tabular Markov Decision ProcessesZhihan Xiong, Ruoqi Shen, Qiwen Cui, Maryam Fazel 等NeurIPS 2022 · 被引用 17 次
- Bilinear Exponential Family of MDPs: Frequentist Regret Bound with Tractable Exploration & PlanningReda Ouhamma, Debabrota Basu, Odalric MaillardAAAI 2023 · 被引用 14 次
