Efficient Online Reinforcement Learning for Diffusion Policy
Haitong Ma, Tianyi Chen, Kai Wang, Na Li, Bo Dai
摘要
Diffusion policies have achieved superior performance in imitation learning and offline reinforcement learning (RL) due to their rich expressiveness. However, the conventional diffusion training procedure requires samples from target distribution, which is impossible in online RL since we cannot sample from the optimal policy. Backpropagating policy gradient through the diffusion process incurs huge computational costs and instability, thus being expensive and not scalable. To enable efficient training of diffusion policies in online RL, we generalize the conventional denoising score matching by reweighting the loss function. The resulting Reweighted Score Matching (RSM) preserves the optimal solution and low computational cost of denoising score matching, while eliminating the need to sample from the target distribution and allowing learning to optimize value functions. We introduce two tractable reweighted loss functions to solve two commonly used policy optimization problems, policy mirror descent and max-entropy policy, resulting in two practical algorithms named Diffusion Policy Mirror Descent (DPMD) and Soft Diffusion Actor-Critic (SDAC). We conducted comprehensive comparisons on MuJoCo benchmarks. The empirical results show that the proposed algorithms outperform recent diffusion-policy online RLs on most tasks, and the DPMD improves more than 120% over soft actor-critic on Humanoid and Ant.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Q-Learning with Adjoint MatchingQiyang Li, Sergey LevineICLR 2026 · 被引用 36 次
- Multi-agent Coordination via Flow MatchingDongsu Lee, Daehee Lee, Amy ZhangICLR 2026 · 被引用 9 次
- Scalable Exploration for High-Dimensional Continuous Control via Value-Guided FlowYunyue Wei, Chenhui Zuo, Yanan SuiICLR 2026 · 被引用 8 次
- Dichotomous Diffusion Policy OptimizationRuiming Liang, Yinan Zheng, Kexin Zheng, Tianyi Tan 等ICLR 2026 · 被引用 7 次
- Relative Entropy Pathwise Policy OptimizationClaas Voelcker, Axel Brunnbauer, Marcel Hussing, Michal Nauman 等ICLR 2026 · 被引用 6 次
它引用的顶会 Paper18
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar 等ICLR 2021 · 被引用 1,270 次
相关 Paper
- Maximum Entropy Reinforcement Learning with Diffusion PolicyXiaoyi Dong, Jian Cheng, Xi Sheryl ZhangICML 2025
- Learning a Diffusion Model Policy from Rewards via Q-Score MatchingMichael Psenka, Alejandro Escontrela, Pieter Abbeel, Yi MaICML 2024 · 被引用 90 次
- EXPO: Stable Reinforcement Learning with Expressive PoliciesPerry Dong, Qiyang Li, Dorsa Sadigh, Chelsea FinnICLR 2026 · 被引用 35 次
- Diffusion Actor-Critic with Entropy RegulatorYinuo Wang, Likun Wang, Yuxuan Jiang, Wenjun Zou 等NeurIPS 2024 · 被引用 105 次
- Mean Flow Policy OptimizationXiaoyi Dong, Xi Zhang, Jian ChengICML 2026
