Reining Generalization in Offline Reinforcement Learning via Representation Distinction
Yi Ma, Hongyao Tang, Dong Li, Zhaopeng Meng
Abstract
Offline Reinforcement Learning (RL) aims to address the challenge of distribution shift between the dataset and the learned policy, where the value of out-of-distribution (OOD) data may be erroneously estimated due to overgeneralization. It has been observed that a considerable portion of the benefits derived from the conservative terms designed by existing offline RL approaches originates from their impact on the learned representation. This observation prompts us to scrutinize the learning dynamics of offline RL, formalize the process of generalization, and delve into the prevalent overgeneralization issue in offline RL. We then investigate the potential to rein the generalization from the representation perspective to enhance offline RL. Finally, we present Representation Distinction (RD), an innovative plug-in method for improving offline RL algorithm performance by explicitly differentiating between the representations of in-sample and OOD state-action pairs generated by the learning policy. Considering scenarios in which the learning policy mirrors the behavioral policy and similar samples may be erroneously distinguished, we suggest a dynamic adjustment mechanism for RD based on an OOD data generator to prevent data representation collapse and further enhance policy performance. We demonstrate the efficacy of our approach by applying RD to designed backbone algorithms and widely-used offline RL algorithms. The proposed RD method significantly improves their performance across various continuous control tasks on D4RL datasets, surpassing several state-of-the-art offline RL algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 22eee410-1d2f-46f5-ae50-630190106f43Cited by top-tier papers10
- No Representation, No Trust: Connecting Representation, Collapse, and Trust Issues in PPOSkander Moalla, Andrea Miele, Daniil Pyatko, Razvan Pascanu et al.NeurIPS 2024 · 35 citations
- Doubly Mild Generalization for Offline Reinforcement LearningYixiu Mao, Qi Wang, Yun Qu, Yuhang Jiang et al.NeurIPS 2024 · 30 citations
- Improving Deep Reinforcement Learning by Reducing the Chain Effect of Value and Policy ChurnHongyao Tang, Glen BersethNeurIPS 2024 · 23 citations
- Squeeze the Soaked Sponge: Efficient Off-policy RFT for Large Language ModelJing Liang, Jinyi Liu, Yi Ma, Hongyao Tang et al.ICLR 2026 · 10 citations
- Adaptive Neighborhood-Constrained Q Learning for Offline Reinforcement LearningYixiu Mao, Yun Qu, Qi (Cheems) Wang, Xiangyang JiNeurIPS 2025 · 3 citations
Builds on30
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
Related papers
- Dynamic Uncertainty Estimation for Offline Reinforcement LearningJiesheng Wang, Lin Li, Wei Wei, Yujia Zhang et al.AAAI 2025 · 2 citations
- Mildly Conservative Q-Learning for Offline Reinforcement LearningJiafei Lyu, Xiaoteng Ma, Xiu Li, Zongqing LuNeurIPS 2022 · 173 citations
- Regularizing a Model-based Policy Stationary Distribution to Stabilize Offline Reinforcement LearningShentao Yang, Yihao Feng, Shujian Zhang, Mingyuan ZhouICML 2022 · 14 citations
- Enhancing Diffusion Policies with Distribution-Matching Generator in Offline Reinforcement LearningXuemin Hu, Shen Li, Yingfen Xu, Bo Tang et al.AAAI 2026 · 1 citation
- Compositional Conservatism: A Transductive Approach in Offline Reinforcement LearningYeda Song, Dongwook Lee, Gunhee KimICLR 2024 · 1 citation
