PsiPhi-Learning: Reinforcement Learning with Demonstrations using Successor Features and Inverse Temporal Difference Learning
Angelos Filos, Clare Lyle, Yarin Gal, Sergey Levine, Natasha Jaques, Gregory Farquhar
Abstract
We study reinforcement learning (RL) with no-reward demonstrations, a setting in which an RL agent has access to additional data from the interaction of other agents with the same environment. However, it has no access to the rewards or goals of these agents, and their objectives and levels of expertise may vary widely. These assumptions are common in multi-agent settings, such as autonomous driving. To effectively use this data, we turn to the framework of successor features. This allows us to disentangle shared features and dynamics of the environment from agent-specific rewards and policies. We propose a multi-task inverse reinforcement learning (IRL) algorithm, called inverse temporal difference learning (ITD), that learns shared state features, alongside per-agent successor features and preference vectors, purely from demonstrations without reward labels. We further show how to seamlessly integrate ITD with learning from online environment interactions, arriving at a novel algorithm for reinforcement learning with demonstrations, called -learning (pronounced `Sci-Fi'). We provide empirical evidence for the effectiveness of -learning as a method for improving RL, IRL, imitation, and few-shot transfer, and derive worst-case bounds for its performance in zero-shot transfer to new tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1db471c5-fbde-449d-9ccf-098b3241a372Cited by top-tier papers6
- Self-Supervised Reinforcement Learning that Transfers using Random FeaturesBoyuan Chen, Chuning Zhu, Pulkit Agrawal, Kaiqing Zhang et al.NeurIPS 2023 · 16 citations
- Learning Successor Features the Simple WayRaymond Chua, Arna Ghosh, Christos Kaplanis, Blake A. Richards et al.NeurIPS 2024 · 14 citations
- Combining Behaviors with the Successor Features KeyboardWilka Carvalho, Andre Saraiva, Angelos Filos, Andrew K. Lampinen et al.NeurIPS 2023 · 13 citations
- Composing Task Knowledge With Modular Successor Feature ApproximatorsWilka Carvalho, Angelos Filos, Richard L. Lewis, Honglak Lee et al.ICLR 2023 · 2 citations
- Learning from All VehiclesDian Chen, Philipp KrähenbühlCVPR 2022
Builds on6
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 685 citations
- Can Autonomous Vehicles Identify, Recover From, and Adapt to Distribution Shifts?Angelos Filos, Panagiotis Tigas, Rowan McAllister, Nicholas Rhinehart et al.ICML 2020 · 225 citations
- Count-Based Exploration with the Successor RepresentationMarlos C. Machado, Marc G. Bellemare, Michael BowlingAAAI 2020 · 206 citations
- Fast Task Inference with Variational Intrinsic Successor FeaturesSteven Hansen, Will Dabney, André Barreto, David Warde-Farley et al.ICLR 2020 · 176 citations
- Deep Imitative Models for Flexible Inference, Planning, and ControlNicholas Rhinehart, Rowan McAllister, Sergey LevineICLR 2020 · 159 citations
Related papers
- Non-Adversarial Inverse Reinforcement Learning via Successor Feature MatchingArnav Kumar Jain, Harley Wiltzer, Jesse Farebrother, Irina Rish et al.ICLR 2025
- Consistent Zero-Shot Imitation with Contrastive Goal InferenceKathryn Wantlin, Chongyi Zheng, Benjamin EysenbachICML 2026 · 1 citation
- Is Inverse Reinforcement Learning Harder than Standard Reinforcement Learning? A Theoretical PerspectiveLei Zhao, Mengdi Wang, Yu BaiICML 2024 · 3 citations
- Learning Shared Safety Constraints from Multi-task DemonstrationsKonwoo Kim, Gokul Swamy, Zuxin Liu, Ding Zhao et al.NeurIPS 2023 · 31 citations
- Offline Inverse RL: New Solution Concepts and Provably Efficient AlgorithmsFilippo Lazzati, Mirco Mutti, Alberto Maria MetelliICML 2024 · 8 citations
