Off-Policy Evaluation via the Regularized Lagrangian
Mengjiao Yang, Ofir Nachum, Bo Dai, Lihong Li, Dale Schuurmans
Abstract
The recently proposed distribution correction estimation (DICE) family of estimators has advanced the state of the art in off-policy evaluation from behavior-agnostic data. While these estimators all perform some form of stationary distribution correction, they arise from different derivations and objective functions. In this paper, we unify these estimators as regularized Lagrangians of the same linear program. The unification allows us to expand the space of DICE estimators to new alternatives that demonstrate improved performance. More importantly, by analyzing the expanded space of estimators both mathematically and empirically we find that dual solutions offer greater flexibility in navigating the tradeoff between optimization stability and estimation bias, and generally provide superior estimates in practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 89224336-9abd-4df3-8d72-0c01207c6f55Cited by top-tier papers66
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 419 citations
- OptiDICE: Offline Policy Optimization via Stationary Distribution Correction EstimationJongmin Lee, Wonseok Jeon, Byung-Jun Lee, Joelle Pineau et al.ICML 2021 · 137 citations
- Benchmarks for Deep Off-Policy EvaluationJustin Fu, Mohammad Norouzi, Ofir Nachum, George Tucker et al.ICLR 2021 · 112 citations
- Reward is enough for convex MDPsTom Zahavy, Brendan O'Donoghue, Guillaume Desjardins, Satinder SinghNeurIPS 2021 · 96 citations
- COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction EstimationJongmin Lee, Cosmin Paduraru, Daniel J. Mankowitz, Nicolas Heess et al.ICLR 2022 · 84 citations
Builds on5
- Minimax Weight and Q-Function Learning for Off-Policy EvaluationMasatoshi Uehara, Jiawei Huang, Nan JiangICML 2020 · 199 citations
- GenDICE: Generalized Offline Estimation of Stationary ValuesRuiyi Zhang, Bo Dai, Lihong Li, Dale SchuurmansICLR 2020 · 184 citations
- Minimax-Optimal Off-Policy Evaluation with Linear Function ApproximationYaqi Duan, Zeyu Jia, Mengdi WangICML 2020 · 161 citations
- GradientDICE: Rethinking Generalized Offline Estimation of Stationary ValuesShangtong Zhang, Bo Liu, Shimon WhitesonICML 2020 · 107 citations
- Doubly Robust Bias Reduction in Infinite Horizon Off-Policy EstimationZiyang Tang, Yihao Feng, Lihong Li, Dengyong Zhou et al.ICLR 2020 · 72 citations
Related papers
- Relaxed Stationary Distribution Correction Estimation for Improved Offline Policy OptimizationWoosung Kim, Donghyeon Ki, Byung-Jun LeeAAAI 2024 · 4 citations
- Offline Imitation from Observation via Primal Wasserstein State Occupancy MatchingKai Yan, Alexander G. Schwing, Yu-Xiong WangICML 2024 · 3 citations
- Diffusion-DICE: In-Sample Diffusion Guidance for Offline Reinforcement LearningLiyuan Mao, Haoran Xu, Xianyuan Zhan, Weinan Zhang et al.NeurIPS 2024 · 49 citations
- Revisiting Distribution Correction Estimation for Offline Imitation Learning with Suboptimal DatasetQuang Anh PHAM, Tien Mai, Akshat KumarICML 2026
- ODICE: Revealing the Mystery of Distribution Correction Estimation via Orthogonal-gradient UpdateLiyuan Mao, Haoran Xu, Weinan Zhang, Xianyuan ZhanICLR 2024 · 23 citations
