Doubly Regularized Markov Decision Processes for Robust Reinforcement Learning
Yiting He, Zhishuai Liu, Pan Xu
摘要
Empirical successes show that regularization improves the stability and efficiency of reinforcement learning (RL), with applications in robotics and post-training of large language models. Yet, theoretical analyses of regularized Markov decision processes (MDPs) have mostly been confined to the standard RL setting. In this work, we investigate regularized MDPs through the lens of robust RL. We introduce a doubly regularized MDP framework that combines policy and dynamics regularizations, enabling robust policy learning while naturally accommodating continuous action spaces. Within this framework, we develop an optimism-based online algorithm and provide the first finite-sample regret guarantees in both tabular and linear settings. Our results show that algorithms for doubly regularized MDPs are as sample-efficient as well-studied robust MDP algorithms, while additionally benefiting from the flexibility of soft policies. We further design practical algorithmic variants for both settings and demonstrate empirically that our approach efficiently and effectively handles function approximation and exploration in large state-action spaces, achieving robust performance.
We then define the optimal value function V * ,β,η h (s) = sup π∈Π V π,β,η h (s) and the optimal Qfunction Q * ,β,η h (s, a) = sup π∈Π Q π,β,η h (s, a) for all (h, s, a) ∈ [H] × S × A, where Π is the set of all possible policies. Correspondingly, the optimal policy π * = π * h H h=1 is defined as the policy that achieves the optimal value function for all (h, s) ∈ [H] × S, that is,
Under the doubly regularized framework, we establish the dynamic programming principle as follows.
Proposition 3.1. For doubly regularized tabular MDPs, it holds that for any policy π and any (h, s, a) ∈ [H] × S × A,
Doubly Regularized Markov Decision Processes for Robust Reinforcement Learning adopts a value iteration framework. We leverage the robust Bellman optimality equation and incorporate the optimism principle (Abbasi-Yadkori et al., 2011) to estimate the robust Q-functions. For the explicit computation expression, we employ the dual formulation corresponding to each fdivergence setting.
Require: transition regularizer β, policy regularizer Ω η . 1: for k = 1, • • • , K do 2:
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- FLAMBE: Structural Complexity and Representation Learning of Low Rank MDPsAlekh Agarwal, Sham M. Kakade, Akshay Krishnamurthy, Wen SunNeurIPS 2020 · 被引用 271 次
- Maximum Entropy RL (Provably) Solves Some Robust RL ProblemsBenjamin Eysenbach, Sergey LevineICLR 2022 · 被引用 244 次
- Representation Learning for Online and Offline RL in Low-rank MDPsMasatoshi Uehara, Xuezhou Zhang, Wen SunICLR 2022 · 被引用 138 次
- Leverage the Average: an Analysis of KL Regularization in Reinforcement LearningNino Vieillard, Tadashi Kozuno, Bruno Scherrer, Olivier Pietquin 等NeurIPS 2020 · 被引用 106 次
- Twice regularized MDPs and the equivalence between robustness and regularizationEsther Derman, Matthieu Geist, Shie MannorNeurIPS 2021 · 被引用 68 次
相关 Paper
- Robust Offline Reinforcement Learning with Linearly Structured f-Divergence RegularizationCheng Tang, Zhishuai Liu, Pan XuICML 2025
- Double Pessimism is Provably Efficient for Distributionally Robust Offline Reinforcement Learning: Generic Algorithm and Robust Partial CoverageJose H. Blanchet, Miao Lu, Tong Zhang, Han ZhongNeurIPS 2023 · 被引用 58 次
- Robust Reinforcement Learning using Least Squares Policy Iteration with Provable Performance GuaranteesKishan Panaganti Badrinath, Dileep KalathilICML 2021 · 被引用 78 次
- Optimistic Planning by Regularized Dynamic ProgrammingAntoine Moulin, Gergely NeuICML 2023 · 被引用 8 次
- Policy Gradient in Robust MDPs with Global Convergence GuaranteeQiuhao Wang, Chin Pang Ho, Marek PetrikICML 2023 · 被引用 43 次
