Zero-Shot Offline Imitation Learning via Optimal Transport
Thomas Rupf, Marco Bagatella, Nico Gürtler, Jonas Frey, Georg Martius
摘要
Zero-shot imitation learning algorithms hold the promise of reproducing unseen behavior from as little as a single demonstration at test time. Existing practical approaches view the expert demonstration as a sequence of goals, enabling imitation with a high-level goal selector, and a low-level goal-conditioned policy. However, this framework can suffer from myopic behavior: the agent's immediate actions towards achieving individual goals may undermine long-term objectives. We introduce a novel method that mitigates this issue by directly optimizing the occupancy matching objective that is intrinsic to imitation learning. We propose to lift a goal-conditioned value function to a distance between occupancies, which are in turn approximated via a learned world model. The resulting method can learn from offline, suboptimal data, and is capable of non-myopic, zeroshot imitation, as we demonstrate in complex, continuous benchmarks. The code is available at https://github.com/martius-lab/ zilot .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- HOTA: Hamiltonian framework for Optimal Transport AdvectionNazar Buzun, Daniil Shlenskii, Maksim Bobrin, Dmitry V. DylovICLR 2026 · 被引用 2 次
- On Discovering Algorithms for Adversarial Imitation LearningShashank Reddy Chirra, Jayden Teoh, Praveen Paruchuri, Pradeep VarakanthamICLR 2026 · 被引用 1 次
它引用的顶会 Paper26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Behavior Transformers: Cloning modes with one stoneNur Muhammad Shafiullah, Zichen Jeff Cui, Ariuntuya Altanzaya, Lerrel PintoNeurIPS 2022 · 被引用 470 次
- TD-MPC2: Scalable, Robust World Models for Continuous ControlNicklas Hansen, Hao Su, Xiaolong WangICLR 2024 · 被引用 388 次
相关 Paper
- IOSTOM: Offline Imitation Learning from Observations via State Transition Occupancy MatchingQuang Anh Pham, Janaka Chathuranga Brahmanage, Tien Mai, Akshat KumarNeurIPS 2025 · 被引用 2 次
- OSVI-WM: One-Shot Visual Imitation for Unseen Tasks using World-Model-Guided Trajectory GenerationRaktim Gautam Goswami, Prashanth Krishnamurthy, Yann LeCun, Farshad KhorramiNeurIPS 2025 · 被引用 12 次
- Consistent Zero-Shot Imitation with Contrastive Goal InferenceKathryn Wantlin, Chongyi Zheng, Benjamin EysenbachICML 2026 · 被引用 1 次
- Foundation Policies with Hilbert RepresentationsSeohong Park, Tobias Kreiman, Sergey LevineICML 2024 · 被引用 72 次
- Hierarchical Few-Shot Imitation with Skill Transition ModelsKourosh Hakhamaneshi, Ruihan Zhao, Albert Zhan, Pieter Abbeel 等ICLR 2022 · 被引用 51 次
