Reward Fine-Tuning Two-Step Diffusion Models via Learning Differentiable Latent-Space Surrogate Reward
Zhiwei Jia, Yuesong Nan, Huixi Zhao, Gengdai Liu
Abstract
Recent research has shown that fine-tuning diffusion models (DMs) with arbitrary rewards, including nondifferentiable ones, is feasible with reinforcement learning (RL) techniques, enabling flexible model alignment. However, applying existing RL methods to step-distilled DMs is challenging for ultra-fast (≤ 2-step) image generation. Our analysis suggests several limitations of policy-based RL methods such as PPO or DPO toward this goal. Based on the insights, we propose fine-tuning DMs with learned differentiable surrogate rewards. Our method, named LaSRO, learns surrogate reward models in the latent space of SDXL to convert arbitrary rewards into differentiable ones for effective reward gradient guidance. LaSRO leverages pretrained latent DMs for reward modeling and tailors reward optimization for ≤ 2-step image generation with efficient off-policy exploration. LaSRO is effective and stable for improving ultra-fast image generation with different reward objectives, outperforming popular RL methods including DDPO [2] and Diffusion-DPO [71] . We further show LaSRO's connection to value-based RL, providing theoretical insights. See our webpage here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3824c92e-6442-4be8-9d14-cac406758d85Cited by top-tier papers4
- Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion ModelsLuca Eyring, Shyamgopal Karthik, Alexey Dosovitskiy, Nataniel Ruiz et al.NeurIPS 2025 · 36 citations
- Learnable Sparsity for Vision Generative ModelsYang Zhang, Er Jin, Wenzhong Liang, Yanfei Dong et al.ICLR 2026 · 8 citations
- MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiencyNicolas Dufour, Lucas Degeorge, Arijit Ghosh, Vicky Kalogeiton et al.ICML 2026 · 2 citations
- Value Matching: Scalable and Gradient-Free Reward-Guided Flow AdaptationCristian Perez Jensen, Luca Schaufelberger, Riccardo De Santi, Kjell Jorner et al.ICLR 2026
Builds on40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
Related papers
- Reinforcement Learning for Fine-tuning Text-to-Image Diffusion ModelsYing Fan, Olivia Watkins, Yuqing Du, Hao Liu et al.NeurIPS 2023 · 372 citations
- Rethinking DPO-Style Diffusion Aligning FrameworksXun Wu, Shaohan Huang, Lingjie Jiang, Furu WeiICCV 2025 · 4 citations
- Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference OptimizationTao Zhang, Cheng Da, Kun Ding, Huan Yang et al.NeurIPS 2025 · 38 citations
- Using Human Feedback to Fine-tune Diffusion Models without Any Reward ModelKai Yang, Jian Tao, Jiafei Lyu, Chunjiang Ge et al.CVPR 2024 · 34 citations
- Training Diffusion Models with Reinforcement LearningKevin Black, Michael Janner, Yilun Du, Ilya Kostrikov et al.ICLR 2024 · 816 citations
