Learning Pseudometric-based Action Representations for Offline Reinforcement Learning
Pengjie Gu, Mengchen Zhao, Chen Chen, Dong Li, Jianye Hao, Bo An
Abstract
Offline reinforcement learning is a promising approach for practical applications since it does not require interactions with real-world environments. However, existing offline RL methods only work well in environments with continuous or small discrete action spaces. In environments with large and discrete action spaces, such as recommender systems and dialogue systems, the performance of existing methods decreases drastically because they suffer from inaccurate value estimation for a large proportion of out-of-distribution (o.o.d.) actions. While recent works have demonstrated that online RL benefits from incorporating semantic information in action representations, unfortunately, they fail to learn reasonable relative distances between action representations, which is key to offline RL to reduce the influence of o.o.d. actions. This paper proposes an action representation learning framework for offline RL based on a pseudometric, which measures both the behavioral relation and the data-distributional relation between actions. We provide theoretical analysis on the continuity of the expected Q-values and the offline policy improvement using the learned action representations. Experimental results show that our methods significantly improve the performance of two typical offline RL methods in environments with large and discrete action spaces.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 65d204b6-b886-4bb2-935d-ed934d0cb400Cited by top-tier papers7
- Understanding and Addressing the Pitfalls of Bisimulation-based Representations in Offline Reinforcement LearningHongyu Zang, Xin Li, Leiji Zhang, Yang Liu et al.NeurIPS 2023 · 15 citations
- Dynamic Neighborhood Construction for Structured Large Discrete Action SpacesFabian Akkerman, Julius Luy, Wouter van Heeswijk, Maximilian SchifferICLR 2024 · 5 citations
- Controlling Type Confounding in Ad Hoc Teamwork with Instance-wise Teammate Feedback RectificationDong Xing, Pengjie Gu, Qian Zheng, Xinrun Wang et al.ICML 2023 · 4 citations
- Offline RL with Discrete Proxy Representations for Generalizability in POMDPsPengjie Gu, Xinyu Cai, Dong Xing, Xinrun Wang et al.NeurIPS 2023 · 1 citation
- Compositional Conservatism: A Transductive Approach in Offline Reinforcement LearningYeda Song, Dongwook Lee, Gunhee KimICLR 2024 · 1 citation
Builds on11
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Offline Reinforcement Learning with Fisher Divergence Critic RegularizationIlya Kostrikov, Rob Fergus, Jonathan Tompson, Ofir NachumICML 2021 · 350 citations
- Deployment-Efficient Reinforcement Learning via Model-Based Offline OptimizationTatsuya Matsushima, Hiroki Furuta, Yutaka Matsuo, Ofir Nachum et al.ICLR 2021 · 166 citations
- Parrot: Data-Driven Behavioral Priors for Reinforcement LearningAvi Singh, Huihan Liu, Gaoyue Zhou, Albert Yu et al.ICLR 2021 · 161 citations
- Representation Matters: Offline Pretraining for Sequential Decision MakingMengjiao Yang, Ofir NachumICML 2021 · 126 citations
Related papers
- Mildly Conservative Q-Learning for Offline Reinforcement LearningJiafei Lyu, Xiaoteng Ma, Xiu Li, Zongqing LuNeurIPS 2022 · 173 citations
- Behavior Proximal Policy OptimizationZifeng Zhuang, Kun Lei, Jinxin Liu, Donglin Wang et al.ICLR 2023 · 8 citations
- Reining Generalization in Offline Reinforcement Learning via Representation DistinctionYi Ma, Hongyao Tang, Dong Li, Zhaopeng MengNeurIPS 2023 · 19 citations
- Robust Task Representations for Offline Meta-Reinforcement Learning via Contrastive LearningHaoqi Yuan, Zongqing LuICML 2022 · 53 citations
- Offline Reinforcement Learning with Pseudometric LearningRobert Dadashi, Shideh Rezaeifar, Nino Vieillard, Léonard Hussenot et al.ICML 2021 · 44 citations
