Consistent Zero-Shot Imitation with Contrastive Goal Inference
Kathryn Wantlin, Chongyi Zheng, Benjamin Eysenbach
Abstract
Zero-shot imitation learning requires an agent to reproduce expert behavior from a single demonstration without additional environment interaction or gradient updates at test time. We introduce Contrastive Inverse Reinforcement Learning (CIRL), a self-supervised framework for pre-training zero-shot imitation agents. Our methods rests on a key observation that many useful tasks can be summarized by a single goal state. We can thus convert the multi-task inverse RL problem into a more tractable goal-inference problem, and utilize state-of-the-art goal-conditioned RL methods to recover a policy that reaches the goal. During pre-training, CIRL jointly employs three components to learn without any rewards or demonstrations: (1) a variant of contrastive RL designed to learn maximum-entropy goal-conditioned policies, (2) an automatic goal proposal mechanism (GoalKDE) that drives exploration, and (3) a mean-field variational model that performs amortized goal inference from trajectories. We prove that this procedure consistently recovers the demonstrator's intent by accounting for the relative difficulty of reaching different states and show how structurally similar prior work may otherwise fail to infer the correct reward. Experiments on goal-conditioned and standard reward-maximizing control tasks show that CIRL outperforms prior zero-shot imitation methods, supporting the expressiveness of goals as a compact summary of behavior.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 707b86de-0fc0-44dc-abc4-75c399565952Builds on22
- Contrastive Learning as Goal-Conditioned Reinforcement LearningBenjamin Eysenbach, Tianjun Zhang, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2022 · 331 citations
- Skew-Fit: State-Covering Self-Supervised Reinforcement LearningVitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair et al.ICML 2020 · 303 citations
- Explore, Discover and Learn: Unsupervised Discovery of State-Covering SkillsVictor Campos, Alexander Trott, Caiming Xiong, Richard Socher et al.ICML 2020 · 178 citations
- Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement LearningSilviu Pitis, Harris Chan, Stephen Zhao, Bradly C. Stadie et al.ICML 2020 · 145 citations
- Learning One Representation to Optimize All RewardsAhmed Touati, Yann OllivierNeurIPS 2021 · 140 citations
Related papers
- Multi-Modal Inverse Constrained Reinforcement Learning from a Mixture of DemonstrationsGuanren Qiao, Guiliang Liu, Pascal Poupart, Zhiqiang XuNeurIPS 2023 · 28 citations
- PsiPhi-Learning: Reinforcement Learning with Demonstrations using Successor Features and Inverse Temporal Difference LearningAngelos Filos, Clare Lyle, Yarin Gal, Sergey Levine et al.ICML 2021 · 32 citations
- RLZero: Direct Policy Inference from Language Without In-Domain SupervisionHarshit Sikchi, Siddhant Agarwal, Pranaya Jajoo, Samyak Parajuli et al.NeurIPS 2025 · 8 citations
- Demonstration-Conditioned Reinforcement Learning for Few-Shot ImitationChristopher R. Dance, Julien Perez, Théo CachetICML 2021 · 17 citations
- Stabilizing Contrastive RL: Techniques for Robotic Goal Reaching from Offline DataChongyi Zheng, Benjamin Eysenbach, Homer Rich Walke, Patrick Yin et al.ICLR 2024 · 16 citations
