Lune

ICML2025

Strategic Planning: A Top-Down Approach to Option Generation

Max Ruiz Luyten, Antonin Berthon, Mihaela van der Schaar

2025Year

Abstract

Real-world human decision making often relies on strategic planning, where high-level goals guide the formulation of sub-goals and subsequent actions, as evidenced by domains such as healthcare, business, and urban policy. Despite notable successes in controlled settings, conventional reinforcement learning (RL) follows a bottom-up framework, which can struggle to adapt to realworld complexities such as sparse rewards and limited exploration budgets. While methods like hierarchical RL and environment shaping provide partial solutions, they frequently rely on either ad hoc designs (e.g. choosing the set of highlevel actions) or purely data-driven discovery of high-level actions that still requires significant exploration. In this paper, we introduce a top-down RL framework that explicitly leverages humaninspired strategy to reduce sample complexity, guide exploration, and enable high-level decision making. We first formalize the Strategy Problem, which frames policy generation as finding distributions over policies that balance specificity and value. Building on this definition, we propose the Strategist agent-an iterative framework that leverages large language models to synthesize domain knowledge into a structured representation of actionable strategies and sub-goals. We further develop a reward shaping methodology that translates these strategies expressed in natural language into quantitative feedback for RL methods. Empirically, we demonstrate that our framework significantly enhances the performance of different underlying RL algorithms, leading to faster convergence and the discovery of more complex behaviors. Taken together, our findings highlight that top-down strategic exploration opens new avenues to improve RL in real-world decision problems.