A First-Occupancy Representation for Reinforcement Learning
Ted Moskovitz, Spencer R. Wilson, Maneesh Sahani
Abstract
Both animals and artificial agents benefit from state representations that support rapid transfer of learning across tasks and which enable them to efficiently traverse their environments to reach rewarding states. The successor representation (SR), which measures the expected cumulative, discounted state occupancy under a fixed policy, enables efficient transfer to different reward structures in an otherwise constant Markovian environment and has been hypothesized to underlie aspects of biological behavior and neural activity. However, in the real world, rewards may move or only be available for consumption once, may shift location, or agents may simply aim to reach goal states as rapidly as possible without the constraint of artificially imposed task horizons. In such cases, the most behaviorally-relevant representation would carry information about when the agent was likely to first reach states of interest, rather than how often it should expect to visit them over a potentially infinite time span. To reflect such demands, we introduce the firstoccupancy representation (FR), which measures the expected temporal discount to the first time a state is accessed. We demonstrate that the FR facilitates exploration, the selection of efficient paths to desired states, allows the agent, under certain conditions, to plan provably optimal trajectories defined by a sequence of subgoals, and induces similar behavior to animals avoiding threatening stimuli. This dichotomy has motivated a search for intermediate models which cache information about environmental structure, and so enable efficient but flexible planning. One such approach, based on the successor representation (SR) (Dayan, 1993) , has been the subject of recent interest in the context of both biological (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Confronting Reward Model Overoptimization with Constrained RLHFTed Moskovitz, Aaditya K. Singh, DJ Strouse, Tuomas Sandholm et al.ICLR 2024 · 89 citations
- Minimum Description Length ControlTed Moskovitz, Ta-Chu Kao, Maneesh Sahani, Matt M. BotvinickICLR 2023 · 76 citations
- Successor-Predecessor Intrinsic ExplorationChangmin Yu, Neil Burgess, Maneesh Sahani, Samuel J. GershmanNeurIPS 2023 · 12 citations
- Reward-Aware Proto-Representations in Reinforcement LearningHon Tik Tse, Siddarth Chandrasekar, Marlos C. MachadoNeurIPS 2025 · 6 citations
- A State Representation for Diminishing RewardsTed Moskovitz, Samo Hromadka, Ahmed Touati, Diana Borsa et al.NeurIPS 2023 · 4 citations
Builds on8
- Dynamics-Aware Unsupervised Discovery of SkillsArchit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar et al.ICLR 2020 · 475 citations
- Count-Based Exploration with the Successor RepresentationMarlos C. Machado, Marc G. Bellemare, Michael BowlingAAAI 2020 · 206 citations
- Fast Task Inference with Variational Intrinsic Successor FeaturesSteven Hansen, Will Dabney, André Barreto, David Warde-Farley et al.ICLR 2020 · 176 citations
- APS: Active Pretraining with Successor FeaturesHao Liu, Pieter AbbeelICML 2021 · 147 citations
- Reinforcement Learning with Non-Markovian RewardsMaor Gaon, Ronen I. BrafmanAAAI 2020 · 96 citations
Related papers
- Contextual Pre-planning on Reward Machine Abstractions for Enhanced Transfer in Deep Reinforcement LearningGuy Azran, Mohamad H. Danesh, Stefano V. Albrecht, Sarah KerenAAAI 2024 · 2 citations
- Hierarchical Successor Representation for Robust TransferChangmin Yu, Máté LengyelICML 2026
- Maximum State Entropy Exploration using Predecessor and Successor RepresentationsArnav Kumar Jain, Lucas Lehnert, Irina Rish, Glen BersethNeurIPS 2023 · 27 citations
- Learning Subgoal Representations with Slow DynamicsSiyuan Li, Lulu Zheng, Jianhao Wang, Chongjie ZhangICLR 2021 · 48 citations
- Successor Feature Sets: Generalizing Successor Representations Across PoliciesKianté Brantley, Soroush Mehri, Geoffrey J. GordonAAAI 2021 · 11 citations
