Lune

ICML2026Top-tier venue

Doubly Regularized Markov Decision Processes for Robust Reinforcement Learning

Yiting He, Zhishuai Liu, Pan Xu

2026Year
213Citations

Abstract

Empirical successes show that regularization improves the stability and efficiency of reinforcement learning (RL), with applications in robotics and post-training of large language models. Yet, theoretical analyses of regularized Markov decision processes (MDPs) have mostly been confined to the standard RL setting. In this work, we investigate regularized MDPs through the lens of robust RL. We introduce a doubly regularized MDP framework that combines policy and dynamics regularizations, enabling robust policy learning while naturally accommodating continuous action spaces. Within this framework, we develop an optimism-based online algorithm and provide the first finite-sample regret guarantees in both tabular and linear settings. Our results show that algorithms for doubly regularized MDPs are as sample-efficient as well-studied robust MDP algorithms, while additionally benefiting from the flexibility of soft policies. We further design practical algorithmic variants for both settings and demonstrate empirically that our approach efficiently and effectively handles function approximation and exploration in large state-action spaces, achieving robust performance.

We then define the optimal value function V * ,β,η h (s) = sup π∈Π V π,β,η h (s) and the optimal Qfunction Q * ,β,η h (s, a) = sup π∈Π Q π,β,η h (s, a) for all (h, s, a) ∈ [H] × S × A, where Π is the set of all possible policies. Correspondingly, the optimal policy π * = π * h H h=1 is defined as the policy that achieves the optimal value function for all (h, s) ∈ [H] × S, that is,

Under the doubly regularized framework, we establish the dynamic programming principle as follows.

Proposition 3.1. For doubly regularized tabular MDPs, it holds that for any policy π and any (h, s, a) ∈ [H] × S × A,

Doubly Regularized Markov Decision Processes for Robust Reinforcement Learning adopts a value iteration framework. We leverage the robust Bellman optimality equation and incorporate the optimism principle (Abbasi-Yadkori et al., 2011) to estimate the robust Q-functions. For the explicit computation expression, we employ the dual formulation corresponding to each fdivergence setting.

Require: transition regularizer β, policy regularizer Ω η . 1: for k = 1, • • • , K do 2:

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 65d776e5-4af0-4ad7-a7eb-d45324df16d9

Builds on12

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines