Mind the Gap: Offline Policy Optimization for Imperfect Rewards
Jianxiong Li, Xiao Hu, Haoran Xu, Jingjing Liu, Xianyuan Zhan, Qing-Shan Jia, Ya-Qin Zhang
Abstract
Reward function is essential in reinforcement learning (RL), serving as the guiding signal to incentivize agents to solve given tasks, however, is also notoriously difficult to design. In many cases, only imperfect rewards are available, which inflicts substantial performance loss for RL agents. In this study, we propose a unified offline policy optimization approach, RGM (Reward Gap Minimization), which can smartly handle diverse types of imperfect rewards. RGM is formulated as a bi-level optimization problem: the upper layer optimizes a reward correction term that performs visitation distribution matching w.r.t. some expert data; the lower layer solves a pessimistic RL problem with the corrected rewards. By exploiting the duality of the lower layer, we derive a tractable algorithm that enables sampled-based learning without any online interactions. Comprehensive experiments demonstrate that RGM achieves superior performance to existing methods under diverse settings of imperfect rewards. Further, RGM can effectively correct wrong or inconsistent rewards against expert preference and retrieve useful information from biased rewards.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers14
- Diffusion-DICE: In-Sample Diffusion Guidance for Offline Reinforcement LearningLiyuan Mao, Haoran Xu, Xianyuan Zhan, Weinan Zhang et al.NeurIPS 2024 · 49 citations
- Survival Instinct in Offline Reinforcement LearningAnqi Li, Dipendra Misra, Andrey Kolobov, Ching-An ChengNeurIPS 2023 · 26 citations
- ODICE: Revealing the Mystery of Distribution Correction Estimation via Orthogonal-gradient UpdateLiyuan Mao, Haoran Xu, Weinan Zhang, Xianyuan ZhanICLR 2024 · 23 citations
- MAHALO: Unifying Offline Reinforcement Learning and Imitation Learning from ObservationsAnqi Li, Byron Boots, Ching-An ChengICML 2023 · 21 citations
- Uni-RLHF: Universal Platform and Benchmark Suite for Reinforcement Learning with Diverse Human FeedbackYifu Yuan, Jianye Hao, Yi Ma, Zibin Dong et al.ICLR 2024 · 21 citations
Builds on20
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- Offline Reinforcement Learning with Fisher Divergence Critic RegularizationIlya Kostrikov, Rob Fergus, Jonathan Tompson, Ofir NachumICML 2021 · 350 citations
- Learning to Utilize Shaping Rewards: A New Approach of Reward ShapingYujing Hu, Weixun Wang, Hangtian Jia, Yixiang Wang et al.NeurIPS 2020 · 256 citations
Related papers
- Bi-Level Offline Policy Optimization with Limited ExplorationWenzhuo ZhouNeurIPS 2023 · 6 citations
- Model-Based Offline Reinforcement Learning with Pessimism-Modulated Dynamics BeliefKaiyang Guo, Yunfeng Shao, Yanhui GengNeurIPS 2022 · 39 citations
- Offline Model-Based Optimization via Policy-Guided Gradient SearchYassine Chemingui, Aryan Deshwal, Trong Nghia Hoang, Janardhan Rao DoppaAAAI 2024 · 22 citations
- Pareto Policy Pool for Model-based Offline Reinforcement LearningYijun Yang, Jing Jiang, Tianyi Zhou, Jie Ma et al.ICLR 2022 · 25 citations
- Dynamic Uncertainty Estimation for Offline Reinforcement LearningJiesheng Wang, Lin Li, Wei Wei, Yujia Zhang et al.AAAI 2025 · 2 citations
