Lune

ICML2026顶会

Doubly Regularized Markov Decision Processes for Robust Reinforcement Learning

Yiting He, Zhishuai Liu, Pan Xu

出版方
2026年份
213被引次数

摘要

Empirical successes show that regularization improves the stability and efficiency of reinforcement learning (RL), with applications in robotics and post-training of large language models. Yet, theoretical analyses of regularized Markov decision processes (MDPs) have mostly been confined to the standard RL setting. In this work, we investigate regularized MDPs through the lens of robust RL. We introduce a doubly regularized MDP framework that combines policy and dynamics regularizations, enabling robust policy learning while naturally accommodating continuous action spaces. Within this framework, we develop an optimism-based online algorithm and provide the first finite-sample regret guarantees in both tabular and linear settings. Our results show that algorithms for doubly regularized MDPs are as sample-efficient as well-studied robust MDP algorithms, while additionally benefiting from the flexibility of soft policies. We further design practical algorithmic variants for both settings and demonstrate empirically that our approach efficiently and effectively handles function approximation and exploration in large state-action spaces, achieving robust performance.

We then define the optimal value function V * ,β,η h (s) = sup π∈Π V π,β,η h (s) and the optimal Qfunction Q * ,β,η h (s, a) = sup π∈Π Q π,β,η h (s, a) for all (h, s, a) ∈ [H] × S × A, where Π is the set of all possible policies. Correspondingly, the optimal policy π * = π * h H h=1 is defined as the policy that achieves the optimal value function for all (h, s) ∈ [H] × S, that is,

Under the doubly regularized framework, we establish the dynamic programming principle as follows.

Proposition 3.1. For doubly regularized tabular MDPs, it holds that for any policy π and any (h, s, a) ∈ [H] × S × A,

Doubly Regularized Markov Decision Processes for Robust Reinforcement Learning adopts a value iteration framework. We leverage the robust Bellman optimality equation and incorporate the optimism principle (Abbasi-Yadkori et al., 2011) to estimate the robust Q-functions. For the explicit computation expression, we employ the dual formulation corresponding to each fdivergence setting.

Require: transition regularizer β, policy regularizer Ω η . 1: for k = 1, • • • , K do 2:

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper12

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖