The Cost of Commitment in Option-Based Hierarchical RL
Randy Lefebvre, Audrey Durand
Abstract
Empirically, option-based hierarchical reinforcement (HRL) learning often produces longer and more diverse options when a deliberation cost is charged at option boundaries. However, when options are executed for many steps under an approximate dynamics model, small model errors compound along the option, degrading the quality of the resulting plan. In this work, we introduce the commitment loss to formalize the tradeoff between deliberation cost and model error as a function of option duration. We characterize how optimal termination probabilities vary with this tradeoff under two model-error mechanisms. First, the model is learned from finite data via maximum-likelihood estimation, producing statistical error that interacts with option duration. Second, we consider an input-driven setting where an exogenous input is only observed at option boundaries and evolves unobserved between them, creating a drift-induced mismatch between planned and realized dynamics. In both cases, we solve for the optimal termination behavior as a function of deliberation cost and the error scale, clarifying the behavior of some popular HRL algorithms that approach the deliberation cost as a heuristic.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 52fcec75-9ac8-4560-8c05-68575e8e45baBuilds on7
- On the Role of Discount Factor in Offline Reinforcement LearningHao Hu, Yiqin Yang, Qianchuan Zhao, Chongjie ZhangICML 2022 · 26 citations
- Creating Multi-Level Skill Hierarchies in Reinforcement LearningJoshua B. Evans, Özgür SimsekNeurIPS 2023 · 15 citations
- Learning Uncertainty-Aware Temporally-Extended ActionsJoongkyu Lee, Seung Joon Park, Yunhao Tang, Min-hwan OhAAAI 2024 · 3 citations
- Temporally-Extended ε-Greedy ExplorationWill Dabney, Georg Ostrovski, André BarretoICLR 2021 · 2 citations
- On Shallow Planning Under Partial ObservabilityRandy Lefebvre, Audrey DurandAAAI 2025 · 2 citations
Related papers
- Reinforcement Learning with a TerminatorGuy Tennenholtz, Nadav Merlis, Lior Shani, Shie Mannor et al.NeurIPS 2022 · 5 citations
- Maximum Likelihood Reinforcement LearningFahim Tajwar, Guanning Zeng, Yueer Zhou, Yuda Song et al.ICML 2026 · 18 citations
- On Rollouts in Model-Based Reinforcement LearningBernd Frauenknecht, Devdutt Subhasish, Friedrich Solowjow, Sebastian TrimpeICLR 2025 · 1 citation
- Effectively Learning Initiation Sets in Hierarchical Reinforcement LearningAkhil Bagaria, Ben Abbatematteo, Omer Gottesman, Matt Corsaro et al.NeurIPS 2023 · 9 citations
- Context-Specific Representation Abstraction for Deep Option LearningMarwa Abdulhai, Dong-Ki Kim, Matthew Riemer, Miao Liu et al.AAAI 2022 · 14 citations
