Hierarchical Reinforcement Learning with Timed Subgoals
Nico Gürtler, Dieter Büchler, Georg Martius
Abstract
Hierarchical reinforcement learning (HRL) holds great potential for sampleefficient learning on challenging long-horizon tasks. In particular, letting a higher level assign subgoals to a lower level has been shown to enable fast learning on difficult problems. However, such subgoal-based methods have been designed with static reinforcement learning environments in mind and consequently struggle with dynamic elements beyond the immediate control of the agent even though they are ubiquitous in real-world problems. In this paper, we introduce Hierarchical reinforcement learning with Timed Subgoals (HiTS), an HRL algorithm that enables the agent to adapt its timing to a dynamic environment by not only specifying what goal state is to be reached but also when. We discuss how communicating with a lower level in terms of such timed subgoals results in a more stable learning problem for the higher level. Our experiments on a range of standard benchmarks and three new challenging dynamic reinforcement learning environments show that our method is capable of sample-efficient learning where an existing state-of-the-art subgoal-based HRL method fails to learn stable solutions. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Learning Hierarchical World Models with Adaptive Temporal Abstractions from Discrete Latent DynamicsChristian Gumbsch, Noor Sajid, Georg Martius, Martin V. ButzICLR 2024 · 24 citations
- Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement LearningHongjoon Ahn, Heewoong Choi, Jisu Han, Taesup MoonNeurIPS 2025 · 22 citations
- Scalable Option Learning in High-Throughput EnvironmentsMikael Henaff, Scott Fujimoto, Michael Matthews, Michael RabbatICML 2026 · 5 citations
- Policy Compatible Skill Incremental Learning via Lazy Learning InterfaceDaehee Lee, Dongsu Lee, TaeYoon Kwack, Wonje Choi et al.NeurIPS 2025 · 3 citations
- Compositional Planning with Jumpy World ModelsJesse Farebrother, Matteo Pirotta, Andrea Tirinzoni, Marc Bellemare et al.ICML 2026 · 1 citation
Builds on1
Related papers
- State-Conditioned Adversarial Subgoal GenerationVivienne Huiling Wang, Joni Pajarinen, Tinghuai Wang, Joni-Kristian KämäräinenAAAI 2023 · 16 citations
- Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel ApproachUtsav Singh, Souradip Chakraborty, Wesley Suttle, Brian M. Sadler et al.ICLR 2026
- CRISP: Curriculum-Inducing Primitive Informed Subgoal Prediction for Boosting Hierarchical Reinforcement LearningUtsav Singh, Vinay P. NamboodiriAAAI 2026 · 6 citations
- PEAR: Primitive Enabled Adaptive Relabeling for Boosting Hierarchical Reinforcement LearningUtsav Singh, Vinay P. NamboodiriICLR 2025
- Learning Subgoal Representations with Slow DynamicsSiyuan Li, Lulu Zheng, Jianhao Wang, Chongjie ZhangICLR 2021 · 48 citations
