Score Regularized Policy Optimization through Diffusion Behavior
Huayu Chen, Cheng Lu, Zhengyi Wang, Hang Su, Jun Zhu
Abstract
Recent developments in offline reinforcement learning have uncovered the immense potential of diffusion modeling, which excels at representing heterogeneous behavior policies. However, sampling from diffusion policies is considerably slow because it necessitates tens to hundreds of iterative inference steps for one action. To address this issue, we propose to extract an efficient deterministic inference policy from critic models and pretrained diffusion behavior models, leveraging the latter to directly regularize the policy gradient with the behavior distribution's score function during optimization. Our method enjoys powerful generative capabilities of diffusion modeling while completely circumventing the computationally intensive and time-consuming diffusion sampling scheme, both during training and evaluation. Extensive results on D4RL tasks show that our method boosts action sampling speed by more than 25 times compared with various leading diffusion-based methods in locomotion tasks, while still maintaining state-of-the-art performance. Code: https://github.com/thu-ml/SRPO .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fbafede0-ce56-491d-8145-59f8713e098eCited by top-tier papers41
- Diffusion Policies Creating a Trust Region for Offline Reinforcement LearningTianyu Chen, Zhendong Wang, Mingyuan ZhouNeurIPS 2024 · 51 citations
- Q-Learning with Adjoint MatchingQiyang Li, Sergey LevineICLR 2026 · 36 citations
- RDT2: Exploring the Scaling Limit of UMI Data Towards Zero-Shot Cross-Embodiment GeneralizationLIU SONGMING, Bangguo Li, Kai Ma, Lingxuan Wu et al.ICML 2026 · 31 citations
- Scaling Offline RL via Efficient and Expressive Shortcut ModelsNicolas A. Espinosa Dice, Yiyi Zhang, Yiding Chen, Bradley Guo et al.NeurIPS 2025 · 28 citations
- Variational Distillation of Diffusion Policies into Mixture of ExpertsHongyi Zhou, Denis Blessing, Ge Li, Onur Celik et al.NeurIPS 2024 · 19 citations
Builds on19
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 StepsCheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen et al.NeurIPS 2022 · 2,653 citations
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score DistillationZhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao et al.NeurIPS 2023 · 1,498 citations
Related papers
- Offline Reinforcement Learning via High-Fidelity Generative Behavior ModelingHuayu Chen, Cheng Lu, Chengyang Ying, Hang Su et al.ICLR 2023 · 6 citations
- Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement LearningChen-Xiao Gao, Chenyang Wu, Mingjun Cao, Chenjun Xiao et al.ICML 2025
- Diffusion Policies as an Expressive Policy Class for Offline Reinforcement LearningZhendong Wang, Jonathan J. Hunt, Mingyuan ZhouICLR 2023 · 33 citations
- Efficient Diffusion Policies For Offline Reinforcement LearningBingyi Kang, Xiao Ma, Chao Du, Tianyu Pang et al.NeurIPS 2023 · 195 citations
- Consistency Models as a Rich and Efficient Policy Class for Reinforcement LearningZihan Ding, Chi JinICLR 2024 · 73 citations
