Robot Policy Learning with Temporal Optimal Transport Reward
Yuwei Fu, Haichao Zhang, Di Wu, Wei Xu, Benoit Boulet
Abstract
Reward specification is one of the most tricky problems in Reinforcement Learning, which usually requires tedious hand engineering in practice. One promising approach to tackle this challenge is to adopt existing expert video demonstrations for policy learning. Some recent work investigates how to learn robot policies from only a single/few expert video demonstrations. For example, reward labeling via Optimal Transport (OT) has been shown to be an effective strategy to generate a proxy reward by measuring the alignment between the robot trajectory and the expert demonstrations. However, previous work mostly overlooks that the OT reward is invariant to temporal order information, which could bring extra noise to the reward signal. To address this issue, in this paper, we introduce the Temporal Optimal Transport (TemporalOT) reward to incorporate temporal order information for learning a more accurate OT-based proxy reward. Extensive experiments on the Meta-world benchmark tasks validate the efficacy of the proposed method. Code is available at: https://github.com/fuyw/TemporalOT
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dbb7f0fb-515e-42db-8f8d-cef5d2cec24eCited by top-tier papers3
- TimeRewarder: Learning Dense Reward from Passive Videos via Frame-wise Temporal DistanceYuyang Liu, Chuan Wen, Yihang Hu, Dinesh Jayaraman et al.ICML 2026 · 7 citations
- Video-Based Optimal Transport for Feedback-Efficient Offline Preference-Based Reinforcement LearningMinh-Tung Luu, Hwanhee Kim, Younghwan Lee, Chang D. YooICML 2026 · 1 citation
- Imitation Learning from a Single Temporally Misaligned VideoWilliam Huey, Huaxiaoyue Wang, Anne Wu, Yoav Artzi et al.ICML 2025
Builds on23
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement LearningDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICLR 2022 · 457 citations
- LIV: Language-Image Representations and Rewards for Robotic ControlYecheng Jason Ma, Vikash Kumar, Amy Zhang, Osbert Bastani et al.ICML 2023 · 212 citations
- RoboCLIP: One Demonstration is Enough to Learn Robot PoliciesSumedh Sontakke, Jesse Zhang, Sébastien M. R. Arnold, Karl Pertsch et al.NeurIPS 2023 · 182 citations
- State Entropy Maximization with Random Encoders for Efficient ExplorationYounggyo Seo, Lili Chen, Jinwoo Shin, Honglak Lee et al.ICML 2021 · 158 citations
Related papers
- Optimal Transport for Offline Imitation LearningYicheng Luo, Zhengyao Jiang, Samuel Cohen, Edward Grefenstette et al.ICLR 2023 · 2 citations
- Meta-Imitation Learning by Watching Video DemonstrationsJiayi Li, Tao Lu, Xiaoge Cao, Yinghao Cai et al.ICLR 2022 · 25 citations
- Imitation Learning from Observation with Automatic Discount SchedulingYuyang Liu, Weijun Dong, Yingdong Hu, Chuan Wen et al.ICLR 2024 · 15 citations
- When a Robot is More Capable than a Human: Learning from Constrained DemonstratorsXinhu Li, Ayush Jain, Zhaojing Yang, Yigit Korkmaz et al.ICLR 2026
- Subtask-Aware Visual Reward Learning from Segmented DemonstrationsChangyeon Kim, Minho Heo, Doohyun Lee, Honglak Lee et al.ICLR 2025
