Offline Hierarchical Reinforcement Learning via Inverse Optimization
Carolin Schmidt, Daniele Gammelli, James Harrison, Marco Pavone, Filipe Rodrigues
Abstract
Hierarchical policies enable strong performance in many sequential decision-making problems, such as those with high-dimensional action spaces, those requiring long-horizon planning, and settings with sparse rewards. However, learning hierarchical policies from static offline datasets presents a significant challenge. Crucially, actions taken by higher-level policies may not be directly observable within hierarchical controllers, and the offline dataset might have been generated using a different policy structure, hindering the use of standard offline learning algorithms. In this work, we propose OHIO: a framework for offline reinforcement learning (RL) of hierarchical policies. Our framework leverages knowledge of the policy structure to solve the inverse problem, recovering the unobservable high-level actions that likely generated the observed data under our hierarchical policy. This approach constructs a dataset suitable for off-the-shelf offline training. We demonstrate our framework on robotic and network optimization problems and show that it substantially outperforms end-to-end RL methods and improves robustness. We investigate a variety of instantiations of our framework, both in direct deployment of policies trained offline and when online fine-tuning is performed. Code and data are available at https://ohio-offline-hierarchical-rl.github.io
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c368b9ae-783c-4505-b348-68778f76bfbdCited by top-tier papers3
- Self-Improving Skill Learning for Robust Skill-based Meta-Reinforcement LearningSeungyul Han, Sanghyeon Lee, Sangjun Bae, Yisak ParkICLR 2026 · 5 citations
- Chain-of-Goals Hierarchical Policy for Long-Horizon Offline Goal-Conditioned RLJinwoo Choi, Sang-Hyun Lee, Seung-Woo SeoICML 2026 · 3 citations
- Hierarchical Decision Making with Structured Policies: A Principled Design via Inverse OptimizationYuexuan Wang, Jingyuan Zhou, Kaidi YangICML 2026
Builds on9
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-TuningMitsuhiko Nakamoto, Simon Zhai, Anikait Singh, Max Sobol Mark et al.NeurIPS 2023 · 296 citations
- Reinforcement Learning with Combinatorial Actions: An Application to Vehicle RoutingArthur Delarue, Ross Anderson, Christian TjandraatmadjaNeurIPS 2020 · 127 citations
- Off-Policy Imitation Learning from ObservationsZhuangdi Zhu, Kaixiang Lin, Bo Dai, Jiayu ZhouNeurIPS 2020 · 102 citations
Related papers
- Flow to Control: Offline Reinforcement Learning with Lossless Primitive DiscoveryYiqin Yang, Hao Hu, Wenzhe Li, Siyuan Li et al.AAAI 2023 · 13 citations
- HIQL: Offline Goal-Conditioned RL with Latent States as ActionsSeohong Park, Dibya Ghosh, Benjamin Eysenbach, Sergey LevineNeurIPS 2023 · 173 citations
- Structural Information-based Hierarchical Diffusion for Offline Reinforcement LearningXianghua Zeng, Hao Peng, Yicheng Pan, Angsheng Li et al.NeurIPS 2025 · 4 citations
- Hierarchical Reinforcement Learning by Discovering Intrinsic OptionsJesse Zhang, Haonan Yu, Wei XuICLR 2021 · 97 citations
- Offline RL with Discrete Proxy Representations for Generalizability in POMDPsPengjie Gu, Xinyu Cai, Dong Xing, Xinrun Wang et al.NeurIPS 2023 · 1 citation
