Optimal Transport for Offline Imitation Learning
Yicheng Luo, Zhengyao Jiang, Samuel Cohen, Edward Grefenstette, Marc Peter Deisenroth
Abstract
With the advent of large datasets, offline reinforcement learning (RL) is a promising framework for learning good decision-making policies without the need to interact with the real environment. However, offline RL requires the dataset to be reward-annotated, which presents practical challenges when reward engineering is difficult or when obtaining reward annotations is labor-intensive. In this paper, we introduce Optimal Transport Reward labeling (OTR), an algorithm that assigns rewards to offline trajectories, with a few high-quality demonstrations. OTR's key idea is to use optimal transport to compute an optimal alignment between an unlabeled trajectory in the dataset and an expert demonstration to obtain a similarity measure that can be interpreted as a reward, which can then be used by an offline RL algorithm to learn the policy. OTR is easy to implement and computationally efficient. On D4RL benchmarks, we show that OTR with a single demonstration can consistently match the performance of offline RL with ground-truth rewards 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext de8c9aa8-5edd-4f1e-8fc0-eb75594ee4f9Cited by top-tier papers25
- Fast Imitation via Behavior Foundation ModelsMatteo Pirotta, Andrea Tirinzoni, Ahmed Touati, Alessandro Lazaric et al.ICLR 2024 · 26 citations
- CEIL: Generalized Contextual Imitation LearningJinxin Liu, Li He, Yachen Kang, Zifeng Zhuang et al.NeurIPS 2023 · 23 citations
- SEABO: A Simple Search-Based Method for Offline Imitation LearningJiafei Lyu, Xiaoteng Ma, Le Wan, Runze Liu et al.ICLR 2024 · 17 citations
- What Matters to You? Towards Visual Representation Alignment for Robot LearningThomas Tian, Chenfeng Xu, Masayoshi Tomizuka, Jitendra Malik et al.ICLR 2024 · 17 citations
- Imitation Learning from Observation with Automatic Discount SchedulingYuyang Liu, Weijun Dong, Yingdong Hu, Chuan Wen et al.ICLR 2024 · 15 citations
Builds on10
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- SQIL: Imitation Learning via Reinforcement Learning with Sparse RewardsSiddharth Reddy, Anca D. Dragan, Sergey LevineICLR 2020 · 299 citations
- Imitation Learning via Off-Policy Distribution MatchingIlya Kostrikov, Ofir Nachum, Jonathan TompsonICLR 2020 · 239 citations
Related papers
- SPRINQL: Sub-optimal Demonstrations driven Offline Imitation LearningHuy Hoang, Tien Mai, Pradeep VarakanthamNeurIPS 2024 · 12 citations
- Rethinking Optimal Transport in Offline Reinforcement LearningArip Asadulaev, Rostislav Korst, Aleksandr Korotin, Vage Egiazarian et al.NeurIPS 2024 · 12 citations
- Should I Run Offline Reinforcement Learning or Behavioral Cloning?Aviral Kumar, Joey Hong, Anikait Singh, Sergey LevineICLR 2022 · 84 citations
- Discriminator-Weighted Offline Imitation Learning from Suboptimal DemonstrationsHaoran Xu, Xianyuan Zhan, Honglei Yin, Huiling QinICML 2022 · 105 citations
- Preference Elicitation for Offline Reinforcement LearningAlizée Pace, Bernhard Schölkopf, Gunnar Rätsch, Giorgia RamponiICLR 2025
