Learning Multimodal Behaviors from Scratch with Diffusion Policy Gradient
Steven Li, Rickmer Krohn, Tao Chen, Anurag Ajay, Pulkit Agrawal, Georgia Chalvatzaki
摘要
Deep reinforcement learning (RL) algorithms typically parameterize the policy as a deep network that outputs either a deterministic action or a stochastic one modeled as a Gaussian distribution, hence restricting learning to a single behavioral mode. Meanwhile, diffusion models emerged as a powerful framework for multimodal learning. However, the use of diffusion policies in online RL is hindered by the intractability of policy likelihood approximation, as well as the greedy objective of RL methods that can easily skew the policy to a single mode. This paper presents Deep Diffusion Policy Gradient (DDiffPG), a novel actor-critic algorithm that learns from scratch multimodal policies parameterized as diffusion models while discovering and maintaining versatile behaviors. DDiffPG explores and discovers multiple modes through off-the-shelf unsupervised clustering combined with novelty-based intrinsic motivation. DDiffPG forms a multimodal training batch and utilizes mode-specific Q-learning to mitigate the inherent greediness of the RL objective, ensuring the improvement of the diffusion policy across all modes. Our approach further allows the policy to be conditioned on mode-specific embeddings to explicitly control the learned modes. Empirical studies validate DDiffPG's capability to master multimodal behaviors in complex, high-dimensional continuous control tasks with sparse rewards, also showcasing proof-of-concept dynamic online replanning when navigating mazes with unseen obstacles.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement LearningTonghe Zhang, Chao Yu, Sichang Su, Yu WangNeurIPS 2025 · 被引用 101 次
- Fine-Tuning Discrete Diffusion Models with Policy Gradient MethodsOussama Zekri, Nicolas BoulléNeurIPS 2025 · 被引用 42 次
- Q-Learning with Adjoint MatchingQiyang Li, Sergey LevineICLR 2026 · 被引用 36 次
- SAC Flow: Sample-Efficient Reinforcement Learning of Flow-Based Policies via Velocity-Reparameterized Sequential ModelingYixian Zhang, Shu'ang Yu, Tonghe Zhang, Mo Guang 等ICLR 2026 · 被引用 33 次
- floq: Training Critics via Flow-Matching for Scaling Compute in Value-Based RLBhavya Agrawalla, Michal Nauman, Khush Agrawal, Aviral KumarICLR 2026 · 被引用 22 次
它引用的顶会 Paper27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
相关 Paper
- Diffusion Policies as an Expressive Policy Class for Offline Reinforcement LearningZhendong Wang, Jonathan J. Hunt, Mingyuan ZhouICLR 2023 · 被引用 33 次
- Reparameterized Policy Learning for Multimodal Trajectory OptimizationZhiao Huang, Litian Liang, Zhan Ling, Xuanlin Li 等ICML 2023 · 被引用 21 次
- Diffusion Actor-Critic with Entropy RegulatorYinuo Wang, Likun Wang, Yuxuan Jiang, Wenjun Zou 等NeurIPS 2024 · 被引用 105 次
- Learning a Diffusion Model Policy from Rewards via Q-Score MatchingMichael Psenka, Alejandro Escontrela, Pieter Abbeel, Yi MaICML 2024 · 被引用 90 次
- Diffusion-based Reinforcement Learning via Q-weighted Variational Policy OptimizationShutong Ding, Ke Hu, Zhenhao Zhang, Kan Ren 等NeurIPS 2024 · 被引用 132 次
