Horizon Generalization in Reinforcement Learning
Vivek Myers, Catherine Ji, Benjamin Eysenbach
Abstract
We study goal-conditioned RL through the lens of generalization, but not in the traditional sense of random augmentations and domain randomization. Rather, we aim to learn goal-directed policies that generalize with respect to the horizon: after training to reach nearby goals (which are easy to learn), these policies should succeed in reaching distant goals (which are quite challenging to learn). In the same way that invariance is closely linked with generalization is other areas of machine learning (e.g., normalization layers make a network invariant to scale, and therefore generalize to inputs of varying scales), we show that this notion of horizon generalization is closely linked with invariance to planning: a policy navigating towards a goal will select the same actions as if it were navigating to a waypoint en route to that goal. Thus, such a policy trained to reach nearby goals should succeed at reaching arbitrarily-distant goals. Our theoretical analysis proves that both horizon generalization and planning invariance are possible, under some assumptions. We present new experimental results and recall findings from prior work in support of our theoretical results. Taken together, our results open the door to studying how techniques for invariance and generalization developed in other areas of machine learning might be adapted to achieve this alluring property. Website and code: https://horizon-generalization.github.io * Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Offline Goal-conditioned Reinforcement Learning with Quasimetric RepresentationsVivek Myers, Bill Zheng, Benjamin Eysenbach, Sergey LevineNeurIPS 2025 · 26 citations
- Temporal Representation Alignment: Successor Features Enable Emergent Compositionality in Robot Instruction FollowingVivek Myers, Bill Zheng, Anca D. Dragan, Kuan Fang et al.NeurIPS 2025 · 13 citations
- Flattening Hierarchies with Policy BootstrappingJohn L. Zhou, Jonathan C. KaoNeurIPS 2025 · 8 citations
- Scaling Goal-conditioned Reinforcement Learning with Multistep Quasimetric DistancesBill Zheng, Vivek Myers, Benjamin Eysenbach, Sergey LevineICLR 2026 · 1 citation
- Compositional Transduction with Latent Analogies for Offline Goal-Conditioned Reinforcement LearningJunseok Kim, Dohyeong Kim, Mineui Hong, Songhwai OhICML 2026 · 1 citation
Builds on27
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Offline Reinforcement Learning as One Big Sequence Modeling ProblemMichael Janner, Qiyang Li, Sergey LevineNeurIPS 2021 · 950 citations
- Reinforcement Learning with Augmented DataMichael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto et al.NeurIPS 2020 · 833 citations
- Contrastive Learning as Goal-Conditioned Reinforcement LearningBenjamin Eysenbach, Tianjun Zhang, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2022 · 331 citations
Related papers
- On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon LengthSunghwan Kim, Junhee Cho, Beong-woo Kwak, Taeyoon Kwon et al.ICML 2026 · 3 citations
- Goal Misgeneralization in Deep Reinforcement LearningLauro Langosco di Langosco, Jack Koch, Lee D. Sharkey, Jacob Pfau et al.ICML 2022 · 128 citations
- C-Planning: An Automatic Curriculum for Learning Goal-Reaching TasksTianjun Zhang, Benjamin Eysenbach, Ruslan Salakhutdinov, Sergey Levine et al.ICLR 2022 · 19 citations
- Learning Domain Invariant Representations in Goal-conditioned Block MDPsBeining Han, Chongyi Zheng, Harris Chan, Keiran Paster et al.NeurIPS 2021 · 20 citations
- Unsupervised Visual Attention and Invariance for Reinforcement LearningXudong Wang, Long Lian, Stella X. YuCVPR 2021
