Exploratory Diffusion Model for Unsupervised Reinforcement Learning
Chengyang Ying, Huayu Chen, Xinning Zhou, Zhongkai Hao, Hang Su, Jun Zhu
Abstract
Unsupervised reinforcement learning (URL) aims to pre-train agents by exploring diverse states or skills in reward-free environments, facilitating efficient adaptation to downstream tasks. As the agent cannot access extrinsic rewards during unsupervised exploration, existing methods design intrinsic rewards to model the explored data and encourage further exploration. However, the explored data are always heterogeneous, posing the requirements of powerful representation abilities for both intrinsic reward models and pre-trained policies. In this work, we propose the Exploratory Diffusion Model (ExDM), which leverages the strong expressive ability of diffusion models to fit the explored data, simultaneously boosting exploration and providing an efficient initialization for downstream tasks. Specifically, ExDM can accurately estimate the distribution of collected data in the replay buffer with the diffusion model and introduces the score-based intrinsic reward, encouraging the agent to explore less-visited states. After obtaining the pre-trained policies, ExDM enables rapid adaptation to downstream tasks. In detail, we provide theoretical analyses and practical algorithms for fine-tuning diffusion policies, addressing key challenges such as training instability and computational complexity caused by multi-step sampling. Extensive experiments demonstrate that ExDM outperforms existing SOTA baselines in efficient unsupervised exploration and fast fine-tuning downstream tasks, especially in structurally complicated environments. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 90bf9709-62b6-43f3-8f53-aef3582b720bCited by top-tier papers4
- Latent State-Predictive Exploration for Deep Reinforcement LearningYiming Wang, Kaiyan Zhao, Borong Zhang, Yan Li et al.AAAI 2026 · 1 citation
- Discovering Mixture Skills for Unsupervised Reinforcement LearningNelson Ma, Junyu Xuan, Guangquan Zhang, Jie LuAAAI 2026
- PACT: Self-Evolving Physical Safety Alignment for Diffusion Policies in Embodied ManipulationLingxuan Wu, Zijian Zhu, Lizhong Wang, Chengyang Ying et al.ICML 2026
- Efficient and Uncertainty-Aware Diffusion Framework for Offline-to-Online Reinforcement LearningHa Manh Bui, Metod Jazbec, Eric Nalisnick, Anqi LiuICML 2026
Builds on44
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 StepsCheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen et al.NeurIPS 2022 · 2,653 citations
Related papers
- Unsupervised Zero-Shot Reinforcement Learning via Dual-Value Forward-Backward RepresentationJingbo Sun, Songjun Tu, Qichao Zhang, Haoran Li et al.ICLR 2025
- Aligning Diffusion Behaviors with Q-functions for Efficient Continuous ControlHuayu Chen, Kaiwen Zheng, Hang Su, Jun ZhuNeurIPS 2024 · 13 citations
- Wasserstein Unsupervised Reinforcement LearningShuncheng He, Yuhang Jiang, Hongchang Zhang, Jianzhun Shao et al.AAAI 2022 · 30 citations
- EUCLID: Towards Efficient Unsupervised Reinforcement Learning with Multi-choice Dynamics ModelYifu Yuan, Jianye Hao, Fei Ni, Yao Mu et al.ICLR 2023 · 1 citation
- Unsupervised Reinforcement Learning with Contrastive Intrinsic ControlMichael Laskin, Hao Liu, Xue Bin Peng, Denis Yarats et al.NeurIPS 2022 · 62 citations
