Fast Imitation via Behavior Foundation Models
Matteo Pirotta, Andrea Tirinzoni, Ahmed Touati, Alessandro Lazaric, Yann Ollivier
摘要
Imitation learning (IL) aims at producing agents that can imitate any behavior given a few expert demonstrations. Yet existing approaches require many demonstrations and/or running (online or offline) reinforcement learning (RL) algorithms for each new imitation task. Here we show that recent RL foundation models based on successor measures can imitate any expert behavior almost instantly with just a few demonstrations and no need for RL or fine-tuning, while accommodating several IL principles (behavioral cloning, feature matching, reward-based, and goal-based reductions). In our experiments, imitation via RL foundation models matches, and often surpasses, the performance of SOTA offline IL algorithms, and produces imitation policies from new demonstrations within seconds instead of hours. * Joint first author, alphabetical order. † Joint last author, alphabetical order. 1 The term "Behavior" emphasizes that the model aims at controlling an agent in a dynamical environment. This avoids confusion with widely used foundation models for images, videos, motions, and language. See Yang et al. ( 2023 ) for an extensive review of the latter for decision making.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- BFM-Zero: A Promptable Behavioral Foundation Model for Humanoid Control Using Unsupervised Reinforcement LearningYitang Li, Zhengyi Luo, Tonghe Zhang, Cunxi Dai 等ICLR 2026 · 被引用 63 次
- Zero-Shot Reinforcement Learning from Low Quality DataScott R. Jeen, Tom Bewley, Jonathan M. CullenNeurIPS 2024 · 被引用 24 次
- TD-JEPA: Latent-predictive Representations for Zero-Shot Reinforcement LearningMarco Bagatella, Matteo Pirotta, Ahmed Touati, Alessandro Lazaric 等ICLR 2026 · 被引用 19 次
- A Distributional Analogue to the Successor RepresentationHarley Wiltzer, Jesse Farebrother, Arthur Gretton, Yunhao Tang 等ICML 2024 · 被引用 11 次
- Zero-Shot Adaptation of Behavioral Foundation Models to Unseen DynamicsMaksim Bobrin, Ilya Zisman, Alexander Nikulin, Vladislav Kurenkov 等ICLR 2026 · 被引用 9 次
它引用的顶会 Paper27
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- IQ-Learn: Inverse soft-Q Learning for ImitationDivyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song 等NeurIPS 2021 · 被引用 271 次
- Imitation Learning via Off-Policy Distribution MatchingIlya Kostrikov, Ofir Nachum, Jonathan TompsonICLR 2020 · 被引用 239 次
- ASE: large-scale reusable adversarial skill embeddings for physically simulated charactersXue Bin Peng, Yunrong Guo, Lina Halper, Sergey Levine 等SIGGRAPH 2022 · 被引用 217 次
- Discovering and Achieving Goals via World ModelsRussell Mendonca, Oleh Rybkin, Kostas Daniilidis, Danijar Hafner 等NeurIPS 2021 · 被引用 177 次
相关 Paper
- Imitation Learning by Reinforcement LearningKamil CiosekICLR 2022 · 被引用 22 次
- Proto Successor Measure: Representing the Behavior Space of an RL AgentSiddhant Agarwal, Harshit Sikchi, Peter Stone, Amy ZhangICML 2025
- Zero-Shot Offline Imitation Learning via Optimal TransportThomas Rupf, Marco Bagatella, Nico Gürtler, Jonas Frey 等ICML 2025
- Optimal Transport for Offline Imitation LearningYicheng Luo, Zhengyao Jiang, Samuel Cohen, Edward Grefenstette 等ICLR 2023 · 被引用 2 次
- RoboCLIP: One Demonstration is Enough to Learn Robot PoliciesSumedh Sontakke, Jesse Zhang, Sébastien M. R. Arnold, Karl Pertsch 等NeurIPS 2023 · 被引用 182 次
