Doubly Regularized Markov Decision Processes for Robust Reinforcement Learning
Yiting He, Zhishuai Liu, Pan Xu
Abstract
Empirical successes show that regularization improves the stability and efficiency of reinforcement learning (RL), with applications in robotics and post-training of large language models. Yet, theoretical analyses of regularized Markov decision processes (MDPs) have mostly been confined to the standard RL setting. In this work, we investigate regularized MDPs through the lens of robust RL. We introduce a doubly regularized MDP framework that combines policy and dynamics regularizations, enabling robust policy learning while naturally accommodating continuous action spaces. Within this framework, we develop an optimism-based online algorithm and provide the first finite-sample regret guarantees in both tabular and linear settings. Our results show that algorithms for doubly regularized MDPs are as sample-efficient as well-studied robust MDP algorithms, while additionally benefiting from the flexibility of soft policies. We further design practical algorithmic variants for both settings and demonstrate empirically that our approach efficiently and effectively handles function approximation and exploration in large state-action spaces, achieving robust performance.
We then define the optimal value function V * ,β,η h (s) = sup π∈Π V π,β,η h (s) and the optimal Qfunction Q * ,β,η h (s, a) = sup π∈Π Q π,β,η h (s, a) for all (h, s, a) ∈ [H] × S × A, where Π is the set of all possible policies. Correspondingly, the optimal policy π * = π * h H h=1 is defined as the policy that achieves the optimal value function for all (h, s) ∈ [H] × S, that is,
Under the doubly regularized framework, we establish the dynamic programming principle as follows.
Proposition 3.1. For doubly regularized tabular MDPs, it holds that for any policy π and any (h, s, a) ∈ [H] × S × A,
Doubly Regularized Markov Decision Processes for Robust Reinforcement Learning adopts a value iteration framework. We leverage the robust Bellman optimality equation and incorporate the optimism principle (Abbasi-Yadkori et al., 2011) to estimate the robust Q-functions. For the explicit computation expression, we employ the dual formulation corresponding to each fdivergence setting.
Require: transition regularizer β, policy regularizer Ω η . 1: for k = 1, • • • , K do 2:
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 65d776e5-4af0-4ad7-a7eb-d45324df16d9Builds on12
- FLAMBE: Structural Complexity and Representation Learning of Low Rank MDPsAlekh Agarwal, Sham M. Kakade, Akshay Krishnamurthy, Wen SunNeurIPS 2020 · 271 citations
- Maximum Entropy RL (Provably) Solves Some Robust RL ProblemsBenjamin Eysenbach, Sergey LevineICLR 2022 · 244 citations
- Representation Learning for Online and Offline RL in Low-rank MDPsMasatoshi Uehara, Xuezhou Zhang, Wen SunICLR 2022 · 138 citations
- Leverage the Average: an Analysis of KL Regularization in Reinforcement LearningNino Vieillard, Tadashi Kozuno, Bruno Scherrer, Olivier Pietquin et al.NeurIPS 2020 · 106 citations
- Twice regularized MDPs and the equivalence between robustness and regularizationEsther Derman, Matthieu Geist, Shie MannorNeurIPS 2021 · 68 citations
Related papers
- Robust Offline Reinforcement Learning with Linearly Structured f-Divergence RegularizationCheng Tang, Zhishuai Liu, Pan XuICML 2025
- Double Pessimism is Provably Efficient for Distributionally Robust Offline Reinforcement Learning: Generic Algorithm and Robust Partial CoverageJose H. Blanchet, Miao Lu, Tong Zhang, Han ZhongNeurIPS 2023 · 58 citations
- Robust Reinforcement Learning using Least Squares Policy Iteration with Provable Performance GuaranteesKishan Panaganti Badrinath, Dileep KalathilICML 2021 · 78 citations
- Optimistic Planning by Regularized Dynamic ProgrammingAntoine Moulin, Gergely NeuICML 2023 · 8 citations
- Policy Gradient in Robust MDPs with Global Convergence GuaranteeQiuhao Wang, Chin Pang Ho, Marek PetrikICML 2023 · 43 citations
