Fast Imitation via Behavior Foundation Models
Matteo Pirotta, Andrea Tirinzoni, Ahmed Touati, Alessandro Lazaric, Yann Ollivier
Abstract
Imitation learning (IL) aims at producing agents that can imitate any behavior given a few expert demonstrations. Yet existing approaches require many demonstrations and/or running (online or offline) reinforcement learning (RL) algorithms for each new imitation task. Here we show that recent RL foundation models based on successor measures can imitate any expert behavior almost instantly with just a few demonstrations and no need for RL or fine-tuning, while accommodating several IL principles (behavioral cloning, feature matching, reward-based, and goal-based reductions). In our experiments, imitation via RL foundation models matches, and often surpasses, the performance of SOTA offline IL algorithms, and produces imitation policies from new demonstrations within seconds instead of hours. * Joint first author, alphabetical order. † Joint last author, alphabetical order. 1 The term "Behavior" emphasizes that the model aims at controlling an agent in a dynamical environment. This avoids confusion with widely used foundation models for images, videos, motions, and language. See Yang et al. ( 2023 ) for an extensive review of the latter for decision making.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e85167d-d0ec-4cb5-9ae4-609c3b430924Cited by top-tier papers17
- BFM-Zero: A Promptable Behavioral Foundation Model for Humanoid Control Using Unsupervised Reinforcement LearningYitang Li, Zhengyi Luo, Tonghe Zhang, Cunxi Dai et al.ICLR 2026 · 63 citations
- Zero-Shot Reinforcement Learning from Low Quality DataScott R. Jeen, Tom Bewley, Jonathan M. CullenNeurIPS 2024 · 24 citations
- TD-JEPA: Latent-predictive Representations for Zero-Shot Reinforcement LearningMarco Bagatella, Matteo Pirotta, Ahmed Touati, Alessandro Lazaric et al.ICLR 2026 · 19 citations
- A Distributional Analogue to the Successor RepresentationHarley Wiltzer, Jesse Farebrother, Arthur Gretton, Yunhao Tang et al.ICML 2024 · 11 citations
- Zero-Shot Adaptation of Behavioral Foundation Models to Unseen DynamicsMaksim Bobrin, Ilya Zisman, Alexander Nikulin, Vladislav Kurenkov et al.ICLR 2026 · 9 citations
Builds on27
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- IQ-Learn: Inverse soft-Q Learning for ImitationDivyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song et al.NeurIPS 2021 · 271 citations
- Imitation Learning via Off-Policy Distribution MatchingIlya Kostrikov, Ofir Nachum, Jonathan TompsonICLR 2020 · 239 citations
- ASE: large-scale reusable adversarial skill embeddings for physically simulated charactersXue Bin Peng, Yunrong Guo, Lina Halper, Sergey Levine et al.SIGGRAPH 2022 · 217 citations
- Discovering and Achieving Goals via World ModelsRussell Mendonca, Oleh Rybkin, Kostas Daniilidis, Danijar Hafner et al.NeurIPS 2021 · 177 citations
Related papers
- Imitation Learning by Reinforcement LearningKamil CiosekICLR 2022 · 22 citations
- Proto Successor Measure: Representing the Behavior Space of an RL AgentSiddhant Agarwal, Harshit Sikchi, Peter Stone, Amy ZhangICML 2025
- Zero-Shot Offline Imitation Learning via Optimal TransportThomas Rupf, Marco Bagatella, Nico Gürtler, Jonas Frey et al.ICML 2025
- Optimal Transport for Offline Imitation LearningYicheng Luo, Zhengyao Jiang, Samuel Cohen, Edward Grefenstette et al.ICLR 2023 · 2 citations
- RoboCLIP: One Demonstration is Enough to Learn Robot PoliciesSumedh Sontakke, Jesse Zhang, Sébastien M. R. Arnold, Karl Pertsch et al.NeurIPS 2023 · 182 citations
