Hierarchical Decision Making with Structured Policies: A Principled Design via Inverse Optimization
Yuexuan Wang, Jingyuan Zhou, Kaidi Yang
摘要
Hierarchical decision-making frameworks are pivotal for addressing complex control tasks, enabling agents to decompose intricate problems into manageable subgoals. Despite their promise, existing hierarchical policies face critical limitations: (i) reinforcement learning (RL)-based methods struggle to guarantee strict constraint satisfaction, and (ii) optimal control (OC)-based approaches often rely on myopic and computationally prohibitive formulations. To reconcile these trade-offs, hierarchical RL-OC architectures have emerged as a promising paradigm. However, the formulation of the lower-level optimization within these frameworks remains underexplored, often relying on heuristic or myopic objectives. In this work, we propose a principled framework that systematically integrates upper-level goal abstraction with structured lower-level decision making. We adopt an inverse optimization approach to inform the structure of the lower-level problem from expert demonstrations, ensuring that the objective of the lower-level policy remains aligned with the overall long-term task goal. To validate the approach, our framework is evaluated on distinct decision making tasks: network-based resource allocation and continuous collision avoidance. Empirical results demonstrate that our method consistently outperforms strong baselines based on end-to-end RL, learning-augmented optimal control, and existing hierarchical RL approaches in both efficiency and decision quality.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- A Hierarchical Reinforcement Learning Based Optimization Framework for Large-scale Dynamic Pickup and Delivery ProblemsYi Ma, Xiaotian Hao, Jianye Hao, Jiawen Lu 等NeurIPS 2021 · 被引用 100 次
- Hierarchical Reinforcement Learning for Integrated RecommendationRuobing Xie, Shaoliang Zhang, Rui Wang, Feng Xia 等AAAI 2021 · 被引用 90 次
- Spatial-Temporal Interplay in Human Mobility: A Hierarchical Reinforcement Learning Approach with Hypergraph RepresentationZhaofan Zhang, Yanan Xiao, Lu Jiang, Dingqi Yang 等AAAI 2024 · 被引用 20 次
- Graph Reinforcement Learning for Network Control via Bi-Level OptimizationDaniele Gammelli, James Harrison, Kaidi Yang, Marco Pavone 等ICML 2023 · 被引用 14 次
- Offline Hierarchical Reinforcement Learning via Inverse OptimizationCarolin Schmidt, Daniele Gammelli, James Harrison, Marco Pavone 等ICLR 2025
相关 Paper
- When Demonstrations meet Generative World Models: A Maximum Likelihood Framework for Offline Inverse Reinforcement LearningSiliang Zeng, Chenliang Li, Alfredo García, Mingyi HongNeurIPS 2023 · 被引用 33 次
- State-Conditioned Adversarial Subgoal GenerationVivienne Huiling Wang, Joni Pajarinen, Tinghuai Wang, Joni-Kristian KämäräinenAAAI 2023 · 被引用 16 次
- Provably Efficient Exploration in Inverse Constrained Reinforcement LearningBo Yue, Jian Li, Guiliang LiuICML 2025
- Simplifying Constraint Inference with Inverse Reinforcement LearningAdriana Hugessen, Harley Wiltzer, Glen BersethNeurIPS 2024 · 被引用 6 次
- Learning Shared Safety Constraints from Multi-task DemonstrationsKonwoo Kim, Gokul Swamy, Zuxin Liu, Ding Zhao 等NeurIPS 2023 · 被引用 31 次
