Score as Action: Fine Tuning Diffusion Generative Models by Continuous-time Reinforcement Learning
Hanyang Zhao, Haoxian Chen, Ji Zhang, David D. Yao, Wenpin Tang
Abstract
Reinforcement learning from human feedback (RLHF), which aligns a diffusion model with input prompt, has become a crucial step in building reliable generative AI models. Most works in this area use a discrete-time formulation, which is prone to induced errors, and often not applicable to models with higher-order/black-box solvers. The objective of this study is to develop a disciplined approach to fine-tune diffusion models using continuous-time RL, formulated as a stochastic control problem with a reward function that aligns the end result (terminal state) with input prompt. The key idea is to treat score matching as controls or actions, and thereby making connections to policy optimization and regularization in continuous-time RL. To carry out this idea, we lay out a new policy optimization framework for continuous-time RL, and illustrate its potential in enhancing the value networks design space via leveraging the structural property of diffusion models. We validate the advantages of our method by experiments in downstream tasks of fine-tuning large-scale Text2Image models of Stable Diffusion v1.5.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Flow-GRPO: Training Flow Matching Models via Online RLJie Liu, Gongye Liu, Jiajun Liang, Yangguang Li et al.NeurIPS 2025 · 647 citations
- Taming Preference Mode Collapse via Directional Decoupling Alignment in Diffusion Reinforcement LearningChubin Chen, Sujie Hu, Jiashu Zhu, Meiqi Wu et al.CVPR 2026 · 28 citations
- The Image as Its Own Reward: Reinforcement Learning with Adversarial Reward for Image GenerationWeijia Mao, Hao Chen, Zhenheng Yang, Mike Zheng ShouCVPR 2026 · 10 citations
- PaCo-RL: Advancing Reinforcement Learning for Consistent Image Generation with Pairwise Reward ModelingBowen Ping, Chengyou Jia, Minnan Luo, Changliang Xia et al.CVPR 2026 · 9 citations
- Continuous Q-Score Matching: Diffusion Guided Reinforcement Learning for Continuous-Time ControlChengxiu Hua, Jiawen Gu, Yushun TangNeurIPS 2025 · 5 citations
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Reinforcement Learning for Fine-tuning Text-to-Image Diffusion ModelsYing Fan, Olivia Watkins, Yuqing Du, Hao Liu et al.NeurIPS 2023 · 372 citations
- Using Human Feedback to Fine-tune Diffusion Models without Any Reward ModelKai Yang, Jian Tao, Jiafei Lyu, Chunjiang Ge et al.CVPR 2024 · 34 citations
- Fine-Tuning Discrete Diffusion Models with Policy Gradient MethodsOussama Zekri, Nicolas BoulléNeurIPS 2025 · 42 citations
- Stochastic Control for Fine-tuning Diffusion Models: Optimality, Regularity, and ConvergenceYinbin Han, Meisam Razaviyayn, Renyuan XuICML 2025
- Diffusion Fine-Tuning via Reparameterized Policy Gradient of the Soft Q-FunctionHyeongyu Kang, Jaewoo Lee, Woocheol Shin, Kiyoung Om et al.ICLR 2026 · 5 citations
