Lune

ICML2026Top-tier venue

REG: In-Sample RL via Regularizing the Evaluation Gap

Hanpu Shen, Weining Shen, Roy Fox

2026Year

Abstract

Distribution shift poses a fundamental challenge in offline reinforcement learning, often leading to value overestimation when querying out-ofdistribution actions. We introduce Regularized Evaluation Gap (REG) as a bridge between implicit methods like IQL and explicit conservative methods. We formulate policy evaluation as a robust optimization problem over an ambiguity set of critics and show that IQL's objective can be viewed as an approximate dual solution to this problem. To extract a policy from the learned value function, we propose a practical Orthogonal Policy Gradient (OPG) update. This method regularizes an aggressive, mode-seeking policy gradient by projecting it onto the subspace orthogonal to a stable, in-sample behavior cloning gradient. Extensive D4RL experiments demonstrate that REG matches state-of-the-art performance among both Gaussian methods and diffusion-based approaches without the computational burden of the latter. 1

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 21720606-9d0c-4004-b40e-fd93d5d04340

Builds on19

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines