Cross-domain Imitation from Observations
Dripta S. Raychaudhuri, Sujoy Paul, Jeroen van Baar, Amit K. Roy-Chowdhury
摘要
Imitation learning seeks to circumvent the difficulty in designing proper reward functions for training agents by utilizing expert behavior. With environments modeled as Markov Decision Processes (MDP), most of the existing imitation algorithms are contingent on the availability of expert demonstrations in the same MDP as the one in which a new imitation policy is to be learned. In this paper, we study the problem of how to imitate tasks when there exist discrepancies between the expert and agent MDP. These discrepancies across domains could include differing dynamics, viewpoint, or morphology; we present a novel framework to learn correspondences across such domains. Importantly, in contrast to prior works, we use unpaired and unaligned trajectories containing only states in the expert domain, to learn this correspondence. We utilize a cycle-consistency constraint on both the state space and a domain agnostic latent space to do this. In addition, we enforce consistency on the temporal position of states via a normalized position estimator function, to align the trajectories across the two domains. Once this correspondence is found, we can directly transfer the demonstrations on one domain to the other and use it for imitation. Experiments across a wide variety of challenging domains demonstrate the efficacy of our approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Cross-Domain Imitation Learning via Optimal TransportArnaud Fickinger, Samuel Cohen, Stuart Russell, Brandon AmosICLR 2022 · 被引用 65 次
- Versatile Offline Imitation from Observations and Examples via Regularized State-Occupancy MatchingYecheng Jason Ma, Andrew Shen, Dinesh Jayaraman, Osbert BastaniICML 2022 · 被引用 49 次
- Robust Imitation Learning against Variations in Environment DynamicsJongseong Chae, Seungyul Han, Whiyoung Jung, Myungsik Cho 等ICML 2022 · 被引用 34 次
- Cross-Domain Policy Adaptation by Capturing Representation MismatchJiafei Lyu, Chenjia Bai, Jingwen Yang, Zongqing Lu 等ICML 2024 · 被引用 30 次
- Learn what matters: cross-domain imitation learning with task-relevant embeddingsTim Franzmeyer, Philip H. S. Torr, João F. HenriquesNeurIPS 2022 · 被引用 28 次
它引用的顶会 Paper3
- Domain Adaptive Imitation LearningKuno Kim, Yihong Gu, Jiaming Song, Shengjia Zhao 等ICML 2020 · 被引用 86 次
- State-only Imitation with Transition Dynamics MismatchTanmay Gangwani, Jian PengICLR 2020 · 被引用 56 次
- RL-CycleGAN: Reinforcement Learning Aware Simulation-to-RealKanishka Rao, Chris Harris, Alex Irpan, Sergey Levine 等CVPR 2020
相关 Paper
- Learning Cross-Domain Correspondence for Control with Dynamics Cycle-ConsistencyQiang Zhang, Tete Xiao, Alexei A. Efros, Lerrel Pinto 等ICLR 2021 · 被引用 73 次
- State Alignment-based Imitation LearningFangchen Liu, Zhan Ling, Tongzhou Mu, Hao SuICLR 2020 · 被引用 103 次
- Robust Visual Imitation Learning with Inverse Dynamics RepresentationsSiyuan Li, Xun Wang, Rongchang Zuo, Kewu Sun 等AAAI 2024 · 被引用 8 次
- Invariant Causal Imitation Learning for Generalizable PoliciesIoana Bica, Daniel Jarrett, Mihaela van der SchaarNeurIPS 2021 · 被引用 46 次
- Domain Adaptive Imitation Learning with Visual ObservationSungho Choi, Seungyul Han, Woojun Kim, Jongseong Chae 等NeurIPS 2023 · 被引用 15 次
