Offline Reinforcement Learning with Pseudometric Learning
Robert Dadashi, Shideh Rezaeifar, Nino Vieillard, Léonard Hussenot, Olivier Pietquin, Matthieu Geist
Abstract
Offline Reinforcement Learning methods seek to learn a policy from logged transitions of an environment, without any interaction. In the presence of function approximation, and under the assumption of limited coverage of the state-action space of the environment, it is necessary to enforce the policy to visit state-action pairs close to the support of logged transitions. In this work, we propose an iterative procedure to learn a pseudometric (closely related to bisimulation metrics) from logged transitions, and use it to define this notion of closeness. We show its convergence and extend it to the function approximation setting. We then use this pseudometric to define a new lookup based bonus in an actor-critic algorithm: PLOFF. This bonus encourages the actor to stay close, in terms of the defined pseudometric, to the support of logged transitions. Finally, we evaluate the method on hand manipulation and locomotion tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2cbb1e35-2194-46f7-86f6-52271d08fbdbCited by top-tier papers10
- Offline Reinforcement Learning as One Big Sequence Modeling ProblemMichael Janner, Qiyang Li, Sergey LevineNeurIPS 2021 · 950 citations
- Pessimistic Q-Learning for Offline Reinforcement Learning: Towards Optimal Sample ComplexityLaixi Shi, Gen Li, Yuting Wei, Yuxin Chen et al.ICML 2022 · 110 citations
- Offline Reinforcement Learning with Value-based Episodic MemoryXiaoteng Ma, Yiqin Yang, Hao Hu, Jun Yang et al.ICLR 2022 · 51 citations
- Curriculum Reinforcement Learning using Optimal Transport via Gradual Domain AdaptationPeide Huang, Mengdi Xu, Jiacheng Zhu, Laixi Shi et al.NeurIPS 2022 · 44 citations
- Is Conditional Generative Modeling all you need for Decision Making?Anurag Ajay, Yilun Du, Abhi Gupta, Joshua B. Tenenbaum et al.ICLR 2023 · 31 citations
Builds on7
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 870 citations
- Scalable Methods for Computing State Similarity in Deterministic Markov Decision ProcessesPablo Samuel CastroAAAI 2020 · 171 citations
- Learning Invariant Representations for Reinforcement Learning without ReconstructionAmy Zhang, Rowan Thomas McAllister, Roberto Calandra, Yarin Gal et al.ICLR 2021 · 77 citations
- Primal Wasserstein Imitation LearningRobert Dadashi, Léonard Hussenot, Matthieu Geist, Olivier PietquinICLR 2021 · 41 citations
Related papers
- Learning Pseudometric-based Action Representations for Offline Reinforcement LearningPengjie Gu, Mengchen Zhao, Chen Chen, Dong Li et al.ICML 2022 · 17 citations
- Generalizable Imitation Learning from Observation via Inferring Goal ProximityYoungwoon Lee, Andrew Szot, Shao-Hua Sun, Joseph J. LimNeurIPS 2021 · 64 citations
- STARC: A General Framework For Quantifying Differences Between Reward FunctionsJoar Max Viktor Skalse, Lucy Farnik, Sumeet Ramesh Motwani, Erik Jenner et al.ICLR 2024 · 14 citations
- Behavior Proximal Policy OptimizationZifeng Zhuang, Kun Lei, Jinxin Liu, Donglin Wang et al.ICLR 2023 · 8 citations
- Mutual Information Regularized Offline Reinforcement LearningXiao Ma, Bingyi Kang, Zhongwen Xu, Min Lin et al.NeurIPS 2023 · 14 citations
