Generating Behaviorally Diverse Policies with Latent Diffusion Models
Shashank Hegde, Sumeet Batra, K. R. Zentner, Gaurav S. Sukhatme
Abstract
Recent progress in Quality Diversity Reinforcement Learning (QD-RL) has enabled learning a collection of behaviorally diverse, high performing policies. However, these methods typically involve storing thousands of policies, which results in high space-complexity and poor scaling to additional behaviors. Condensing the archive into a single model while retaining the performance and coverage of the original collection of policies has proved challenging. In this work, we propose using diffusion models to distill the archive into a single generative model over policy parameters. We show that our method achieves a compression ratio of 13x while recovering 98% of the original rewards and 89% of the original coverage. Further, the conditioning mechanism of diffusion models allows for flexibly selecting and sequencing behaviors, including using language. Project website: https://sites.google.com/view/policydiffusion/home . * Equal contribution † Sukhatme holds concurrent appointments as a Professor at USC and as an Amazon Scholar. This paper describes work performed at USC and is not associated with Amazon. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a9ef290-594c-4e27-b744-b5a38059378bCited by top-tier papers6
- AlignDiff: Aligning Diverse Human Preferences via Behavior-Customisable Diffusion ModelZibin Dong, Yifu Yuan, Jianye Hao, Fei Ni et al.ICLR 2024 · 44 citations
- Generative Modeling of Weights: Generalization or Memorization?Boya Zeng, Yida Yin, Zhiqiu Xu, Zhuang LiuCVPR 2026 · 12 citations
- Quality-Diversity with Limited ResourcesRen-Jian Wang, Ke Xue, Cong Guan, Chao QianICML 2024 · 4 citations
- Learning Intractable Multimodal Policies with Reparameterization and Diversity RegularizationZiqi Wang, Jiashun Liu, Ling PanNeurIPS 2025 · 3 citations
- Discount Model Search for Quality Diversity Optimization in High-Dimensional Measure SpacesBryon Tjanaka, Henry Chen, Matthew Christopher Fontaine, Stefanos NikolaidisICLR 2026 · 1 citation
Builds on12
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
Related papers
- Diffusion Policies as an Expressive Policy Class for Offline Reinforcement LearningZhendong Wang, Jonathan J. Hunt, Mingyuan ZhouICLR 2023 · 33 citations
- Learning a Diffusion Model Policy from Rewards via Q-Score MatchingMichael Psenka, Alejandro Escontrela, Pieter Abbeel, Yi MaICML 2024 · 90 citations
- Offline Reinforcement Learning via High-Fidelity Generative Behavior ModelingHuayu Chen, Cheng Lu, Chengyang Ying, Hang Su et al.ICLR 2023 · 6 citations
- DiffuSeq: Sequence to Sequence Text Generation with Diffusion ModelsShansan Gong, Mukai Li, Jiangtao Feng, Zhiyong Wu et al.ICLR 2023 · 94 citations
- Score Regularized Policy Optimization through Diffusion BehaviorHuayu Chen, Cheng Lu, Zhengyi Wang, Hang Su et al.ICLR 2024 · 59 citations
