Generative Planning for Temporally Coordinated Exploration in Reinforcement Learning
Haichao Zhang, Wei Xu, Haonan Yu
Abstract
Standard model-free reinforcement learning algorithms optimize a policy that generates the action to be taken in the current time step in order to maximize expected future return. While flexible, it faces difficulties arising from the inefficient exploration due to its single step nature. In this work, we present Generative Planning method (GPM), which can generate actions not only for the current step, but also for a number of future steps (thus termed as generative planning). This brings several benefits to GPM. Firstly, since GPM is trained by maximizing value, the plans generated from it can be regarded as intentional action sequences for reaching high value regions. GPM can therefore leverage its generated multi-step plans for temporally coordinated exploration towards high value regions, which is potentially more effective than a sequence of actions generated by perturbing each action at single step level, whose consistent movement decays exponentially with the number of exploration steps. Secondly, starting from a crude initial plan generator, GPM can refine it to be adaptive to the task, which, in return, benefits future explorations. This is potentially more effective than commonly used action-repeat strategy, which is non-adaptive in its form of plans. Additionally, since the multi-step plan can be interpreted as the intent of the agent from now to a span of time period into the future, it offers a more informative and intuitive signal for interpretation. Experiments are conducted on several benchmark environments and the results demonstrated its effectiveness compared with several baseline methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cb09114f-1f25-46aa-9a0c-4db34452e4e9Cited by top-tier papers6
- Diffusion Model is an Effective Planner and Data Synthesizer for Multi-Task Reinforcement LearningHaoran He, Chenjia Bai, Kang Xu, Zhuoran Yang et al.NeurIPS 2023 · 165 citations
- FuRL: Visual-Language Models as Fuzzy Rewards for Reinforcement LearningYuwei Fu, Haichao Zhang, Di Wu, Wei Xu et al.ICML 2024 · 31 citations
- Chunking the Critic: A Transformer-based Soft Actor-Critic with N-Step ReturnsDong Tian, Onur Celik, Gerhard NeumannICLR 2026 · 18 citations
- Policy Expansion for Bridging Offline-to-Online Reinforcement LearningHaichao Zhang, Wei Xu, Haonan YuICLR 2023 · 5 citations
- Overcoming Slow Decision Frequencies in Continuous Control: Model-Based Sequence Reinforcement Learning for Model-Free ControlDevdhar Patel, Hava T. SiegelmannICLR 2025
Builds on11
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel et al.ICML 2020 · 489 citations
- Skew-Fit: State-Covering Self-Supervised Reinforcement LearningVitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair et al.ICML 2020 · 303 citations
- Exploring Model-based Planning with Policy NetworksTingwu Wang, Jimmy BaICLR 2020 · 164 citations
- Deep Imitative Models for Flexible Inference, Planning, and ControlNicholas Rhinehart, Rowan McAllister, Sergey LevineICLR 2020 · 159 citations
Related papers
- GenPlan: Generative Sequence Models as Adaptive PlannersAkash Karthikeyan, Yash Vardhan PantAAAI 2025
- Action abstractions for amortized samplingOussama Boussif, Léna Néhale Ezzine, Joseph D. Viviano, Michal Koziarski et al.ICLR 2025
- PhyPlan: Learning to Plan Tasks with Generalizable and Rapid Physical Reasoning for Embodied ManipulationAnkit Kanwar, Hartej Soin, Abhinav Barnawal, Mudit Chopra et al.AAAI 2026
- PlanGAN: Model-based Planning With Sparse Rewards and Multiple GoalsHenry Charlesworth, Giovanni MontanaNeurIPS 2020 · 34 citations
- Interpreting Emergent Planning in Model-Free Reinforcement LearningThomas Bush, Stephen Chung, Usman Anwar, Adrià Garriga-Alonso et al.ICLR 2025
