Local Linearity: the Key for No-regret Reinforcement Learning in Continuous MDPs
Davide Maran, Alberto Maria Metelli, Matteo Papini, Marcello Restelli
摘要
Achieving the no-regret property for Reinforcement Learning (RL) problems in continuous state and action-space environments is one of the major open problems in the field. Existing solutions either work under very specific assumptions or achieve bounds that are vacuous in some regimes. Furthermore, many structural assumptions are known to suffer from a provably unavoidable exponential dependence on the time horizon in the regret, which makes any possible solution unfeasible in practice. In this paper, we identify local linearity as the feature that makes Markov Decision Processes (MDPs) both learnable (sublinear regret) and feasible (regret that is polynomial in ). We define a novel MDP representation class, namely Locally Linearizable MDPs, generalizing other representation classes like Linear MDPs and MDPS with low inherent Belmman error. Then, i) we introduce Cinderella, a no-regret algorithm for this general representation class, and ii) we show that all known learnable and feasible MDP families are representable in this class. We first show that all known feasible MDPs belong to a family that we call Mildly Smooth MDPs. Then, we show how any mildly smooth MDP can be represented as a Locally Linearizable MDP by an appropriate choice of representation. This way, Cinderella is shown to achieve state-of-the-art regret bounds for all previously known (and some new) continuous MDPs for which RL is learnable and feasible.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- A Differential and Pointwise Control Approach to Reinforcement LearningMinh H. Nguyen, Chandrajit BajajNeurIPS 2025 · 被引用 3 次
- Beyond Least Squares: Uniform Approximation and the Hidden Cost of MisspecificationDavide Maran, Csaba SzepesváriNeurIPS 2025 · 被引用 2 次
它引用的顶会 Paper12
- Bellman Eluder Dimension: New Rich Classes of RL Problems, and Sample-Efficient AlgorithmsChi Jin, Qinghua Liu, Sobhan MiryoosefiNeurIPS 2021 · 被引用 264 次
- Learning Near Optimal Policies with Low Inherent Bellman ErrorAndrea Zanette, Alessandro Lazaric, Mykel J. Kochenderfer, Emma BrunskillICML 2020 · 被引用 238 次
- Provably Efficient Reinforcement Learning with Kernel and Neural Function ApproximationsZhuoran Yang, Chi Jin, Zhaoran Wang, Mengdi Wang 等NeurIPS 2020 · 被引用 48 次
- Control Frequency Adaptation via Action Persistence in Batch Reinforcement LearningAlberto Maria Metelli, Flavio Mazzolini, Lorenzo Bisi, Luca Sabbioni 等ICML 2020 · 被引用 43 次
- Metrics and Continuity in Reinforcement LearningCharline Le Lan, Marc G. Bellemare, Pablo Samuel CastroAAAI 2021 · 被引用 40 次
相关 Paper
- No-Regret Reinforcement Learning in Smooth MDPsDavide Maran, Alberto Maria Metelli, Matteo Papini, Marcello RestelliICML 2024 · 被引用 6 次
- Provably Efficient Lifelong Reinforcement Learning with Linear RepresentationSanae Amani, Lin Yang, Ching-An ChengICLR 2023
- Reinforcement Learning in Linear MDPs: Constant Regret and Representation SelectionMatteo Papini, Andrea Tirinzoni, Aldo Pacchiano, Marcello Restelli 等NeurIPS 2021 · 被引用 26 次
- Achieving Constant Regret in Linear Markov Decision ProcessesWeitong Zhang, Zhiyuan Fan, Jiafan He, Quanquan GuNeurIPS 2024 · 被引用 6 次
- RL for Latent MDPs: Regret Guarantees and a Lower BoundJeongyeol Kwon, Yonathan Efroni, Constantine Caramanis, Shie MannorNeurIPS 2021 · 被引用 91 次
