The Cost of Commitment in Option-Based Hierarchical RL
Randy Lefebvre, Audrey Durand
摘要
Empirically, option-based hierarchical reinforcement (HRL) learning often produces longer and more diverse options when a deliberation cost is charged at option boundaries. However, when options are executed for many steps under an approximate dynamics model, small model errors compound along the option, degrading the quality of the resulting plan. In this work, we introduce the commitment loss to formalize the tradeoff between deliberation cost and model error as a function of option duration. We characterize how optimal termination probabilities vary with this tradeoff under two model-error mechanisms. First, the model is learned from finite data via maximum-likelihood estimation, producing statistical error that interacts with option duration. Second, we consider an input-driven setting where an exogenous input is only observed at option boundaries and evolves unobserved between them, creating a drift-induced mismatch between planned and realized dynamics. In both cases, we solve for the optimal termination behavior as a function of deliberation cost and the error scale, clarifying the behavior of some popular HRL algorithms that approach the deliberation cost as a heuristic.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- On the Role of Discount Factor in Offline Reinforcement LearningHao Hu, Yiqin Yang, Qianchuan Zhao, Chongjie ZhangICML 2022 · 被引用 26 次
- Creating Multi-Level Skill Hierarchies in Reinforcement LearningJoshua B. Evans, Özgür SimsekNeurIPS 2023 · 被引用 15 次
- Learning Uncertainty-Aware Temporally-Extended ActionsJoongkyu Lee, Seung Joon Park, Yunhao Tang, Min-hwan OhAAAI 2024 · 被引用 3 次
- Temporally-Extended ε-Greedy ExplorationWill Dabney, Georg Ostrovski, André BarretoICLR 2021 · 被引用 2 次
- On Shallow Planning Under Partial ObservabilityRandy Lefebvre, Audrey DurandAAAI 2025 · 被引用 2 次
相关 Paper
- Reinforcement Learning with a TerminatorGuy Tennenholtz, Nadav Merlis, Lior Shani, Shie Mannor 等NeurIPS 2022 · 被引用 5 次
- Maximum Likelihood Reinforcement LearningFahim Tajwar, Guanning Zeng, Yueer Zhou, Yuda Song 等ICML 2026 · 被引用 18 次
- On Rollouts in Model-Based Reinforcement LearningBernd Frauenknecht, Devdutt Subhasish, Friedrich Solowjow, Sebastian TrimpeICLR 2025 · 被引用 1 次
- Effectively Learning Initiation Sets in Hierarchical Reinforcement LearningAkhil Bagaria, Ben Abbatematteo, Omer Gottesman, Matt Corsaro 等NeurIPS 2023 · 被引用 9 次
- Context-Specific Representation Abstraction for Deep Option LearningMarwa Abdulhai, Dong-Ki Kim, Matthew Riemer, Miao Liu 等AAAI 2022 · 被引用 14 次
