Relaxed Stationary Distribution Correction Estimation for Improved Offline Policy Optimization
Woosung Kim, Donghyeon Ki, Byung-Jun Lee
摘要
One of the major challenges of offline reinforcement learning (RL) is dealing with distribution shifts that stem from the mismatch between the trained policy and the data collection policy. Stationary distribution correction estimation algorithms (DICE) have addressed this issue by regularizing the policy optimization with f -divergence between the state-action visitation distributions of the data collection policy and the optimized policy. While such regularization naturally integrates to derive an objective to get optimal state-action visitation, such an implicit policy optimization framework has shown limited performance in practice. We observe that the reduced performance is attributed to the biased estimate and the properties of conjugate functions of f -divergence regularization. In this paper, we improve the regularized implicit policy optimization framework by relieving the bias and reshaping the conjugate function by relaxing the constraints. We show that the relaxation adjusts the degree of involvement of the suboptimal samples in optimization, and we derive a new offline RL algorithm that benefits from the relaxed framework, improving from a previous implicit policy optimization algorithm by a large margin.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Prior-Guided Diffusion Planning for Offline Reinforcement LearningDonghyeon Ki, JunHyeok Oh, Seong-Woong Shim, Byung-Jun LeeNeurIPS 2025 · 被引用 16 次
- SEMDICE: Off-policy State Entropy Maximization via Stationary Distribution Correction EstimationJongmin Lee, Meiqi Sun, Pieter AbbeelICLR 2025
它引用的顶会 Paper4
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- OptiDICE: Offline Policy Optimization via Stationary Distribution Correction EstimationJongmin Lee, Wonseok Jeon, Byung-Jun Lee, Joelle Pineau 等ICML 2021 · 被引用 137 次
- Offline RL with No OOD Actions: In-Sample Learning via Implicit Value RegularizationHaoran Xu, Li Jiang, Jianxiong Li, Zhuoran Yang 等ICLR 2023 · 被引用 4 次
相关 Paper
- Offline Imitation Learning with Suboptimal Demonstrations via Relaxed Distribution MatchingLantao Yu, Tianhe Yu, Jiaming Song, Willie Neiswanger 等AAAI 2023 · 被引用 29 次
- ODICE: Revealing the Mystery of Distribution Correction Estimation via Orthogonal-gradient UpdateLiyuan Mao, Haoran Xu, Weinan Zhang, Xianyuan ZhanICLR 2024 · 被引用 23 次
- Offline Imitation from Observation via Primal Wasserstein State Occupancy MatchingKai Yan, Alexander G. Schwing, Yu-Xiong WangICML 2024 · 被引用 3 次
- Revisiting Distribution Correction Estimation for Offline Imitation Learning with Suboptimal DatasetQuang Anh PHAM, Tien Mai, Akshat KumarICML 2026
- Off-Policy Evaluation via the Regularized LagrangianMengjiao Yang, Ofir Nachum, Bo Dai, Lihong Li 等NeurIPS 2020 · 被引用 125 次
