Successor Feature Sets: Generalizing Successor Representations Across Policies
Kianté Brantley, Soroush Mehri, Geoffrey J. Gordon
Abstract
Successor-style representations have many advantages for reinforcement learning: for example, they can help an agent generalize from past experience to new goals, and they have been proposed as explanations of behavioral and neural data from human and animal learners. They also form a natural bridge between model-based and model-free RL methods: like the former they make predictions about future experiences, and like the latter they allow efficient prediction of total discounted rewards. However, successor-style representations are not optimized to generalize across policies: typically, we maintain a limited-length list of policies, and share information among them by representation learning or GPI. Successor-style representations also typically make no provision for gathering information or reasoning about latent variables. To address these limitations, we bring together ideas from predictive state representations, belief space value iteration, successor features, and convex analysis: we develop a new, general successor-style representation, together with a Bellman equation that connects multiple sources of information within this representation, including different latent states, policies, and reward functions. The new representation is highly expressive: for example, it lets us efficiently read off an optimal policy for a new reward function, or a policy that imitates a new demonstration. For this paper, we focus on exact computation of the new representation in small, known environments, since even this restricted setting offers plenty of interesting questions. Our implementation does not scale to large, unknown environments --- nor would we expect it to, since it generalizes POMDP value iteration, which is difficult to scale. However, we believe that future work will allow us to extend our ideas to approximate reasoning in large, unknown environments. We conduct experiments to explore which of the potential barriers to scaling are most pressing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b10a480e-5fb6-4ca1-87e2-cb1c754f6fa6Cited by top-tier papers4
- Learning Successor Features the Simple WayRaymond Chua, Arna Ghosh, Christos Kaplanis, Blake A. Richards et al.NeurIPS 2024 · 14 citations
- π2vec: Policy Representation with Successor FeaturesGianluca Scarpellini, Ksenia Konyushkova, Claudio Fantacci, Thomas Paine et al.ICLR 2024 · 2 citations
- A Generalized Bootstrap Target for Value-Learning, Efficiently Combining Value and Feature PredictionsAnthony GX-Chen, Veronica Chelu, Blake A. Richards, Joelle PineauAAAI 2022 · 1 citation
- When is Transfer Learning Possible?My Phan, Kianté Brantley, Stephanie Milani, Soroush Mehri et al.ICML 2024
Builds on1
Related papers
- Predictive Coding Enhances Meta-RL To Achieve Interpretable Bayes-Optimal Belief Representation Under Partial ObservabilityPo-Chen Kuo, Han Hou, Will Dabney, Edgar Y. WalkerNeurIPS 2025
- Composing Task Knowledge With Modular Successor Feature ApproximatorsWilka Carvalho, Angelos Filos, Richard L. Lewis, Honglak Lee et al.ICLR 2023 · 2 citations
- Value-driven Hindsight ModellingArthur Guez, Fabio Viola, Theophane Weber, Lars Buesing et al.NeurIPS 2020 · 12 citations
- What about Inputting Policy in Value Function: Policy Representation and Policy-Extended Value Function ApproximatorHongyao Tang, Zhaopeng Meng, Jianye Hao, Chen Chen et al.AAAI 2022 · 20 citations
- Bootstrap Latent-Predictive Representations for Multitask Reinforcement LearningZhaohan Daniel Guo, Bernardo Ávila Pires, Bilal Piot, Jean-Bastien Grill et al.ICML 2020 · 153 citations
