Hierarchical Diffusion for Offline Decision Making
Wenhao Li, Xiangfeng Wang, Bo Jin, Hongyuan Zha
摘要
Offline reinforcement learning typically introduces a hierarchical structure to solve the longhorizon problem so as to address its thorny issue of variance accumulation. Problems of deadly triad, limited data and reward sparsity, however, still remain, rendering the design of effective, hierarchical offline RL algorithms for generalpurpose policy learning a formidable challenge. In this paper, we first formulate the problem of offline long-horizon decision-MakIng from the perspective of conditional generative modeling by incorporating goals into the control-as-inference graphic models. A Hierarchical trajectory-level Diffusion probabilistic model is then proposed with classifier-free guidance. HDMI employs a cascade framework that utilizes the rewardconditional goal diffuser for the subgoal discovery and the goal-conditional trajectory diffuser for generating the corresponding action sequence of subgoals. Planning-based subgoal extraction and transformer-based diffusion are employed to deal with the sub-optimal data pollution and long-range subgoal dependencies in the goal diffusion. Numerical experiments verify the advantages of HDMI on long-horizon decision-making compared to SOTA offline RL methods and conditional generative models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper43
- Diffusion-based Reinforcement Learning via Q-weighted Variational Policy OptimizationShutong Ding, Ke Hu, Zhenhao Zhang, Kan Ren 等NeurIPS 2024 · 被引用 132 次
- Simple Hierarchical Planning with DiffusionChang Chen, Fei Deng, Kenji Kawaguchi, Caglar Gulcehre 等ICLR 2024 · 被引用 79 次
- Learning Multimodal Behaviors from Scratch with Diffusion Policy GradientSteven Li, Rickmer Krohn, Tao Chen, Anurag Ajay 等NeurIPS 2024 · 被引用 61 次
- DiffuserLite: Towards Real-time Diffusion PlanningZibin Dong, Jianye Hao, Yifu Yuan, Fei Ni 等NeurIPS 2024 · 被引用 57 次
- Generative Trajectory Stitching through Diffusion CompositionYunhao Luo, Utkarsh A. Mishra, Yilun Du, Danfei XuNeurIPS 2025 · 被引用 48 次
它引用的顶会 Paper38
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 被引用 5,234 次
相关 Paper
- Structural Information-based Hierarchical Diffusion for Offline Reinforcement LearningXianghua Zeng, Hao Peng, Yicheng Pan, Angsheng Li 等NeurIPS 2025 · 被引用 4 次
- Hierarchical Entity-centric Reinforcement Learning with Factored Subgoal DiffusionDan Haramati, Carl Qi, Tal Daniel, Amy Zhang 等ICLR 2026 · 被引用 7 次
- Chain-of-Goals Hierarchical Policy for Long-Horizon Offline Goal-Conditioned RLJinwoo Choi, Sang-Hyun Lee, Seung-Woo SeoICML 2026 · 被引用 3 次
- PlanDQ: Hierarchical Plan Orchestration via D-Conductor and Q-PerformerChang Chen, Junyeob Baek, Fei Deng, Kenji Kawaguchi 等ICML 2024 · 被引用 4 次
- Prior-Guided Diffusion Planning for Offline Reinforcement LearningDonghyeon Ki, JunHyeok Oh, Seong-Woong Shim, Byung-Jun LeeNeurIPS 2025 · 被引用 16 次
