Adding Conditional Control to Diffusion Models with Reinforcement Learning
Yulai Zhao, Masatoshi Uehara, Gabriele Scalia, Sun-Yuan Kung, Tommaso Biancalani, Sergey Levine, Ehsan Hajiramezanali
Abstract
Diffusion models are powerful generative models that allow for precise control over the characteristics of the generated samples. While these diffusion models trained on large datasets have achieved success, there is often a need to introduce additional controls in downstream fine-tuning processes, treating these powerful models as pre-trained diffusion models. This work presents a novel method based on reinforcement learning (RL) to add such controls using an offline dataset comprising inputs and labels. We formulate this task as an RL problem, with the classifier learned from the offline dataset and the KL divergence against pre-trained models serving as the reward functions. Our method, (onditioning pre-rained diffusion models with einforcement earning), produces soft-optimal policies that maximize the abovementioned reward functions. We formally demonstrate that our method enables sampling from the conditional distribution with additional controls during inference. Our RL-based approach offers several advantages over existing methods. Compared to classifier-free guidance, it improves sample efficiency and can greatly simplify dataset construction by leveraging conditional independence between the inputs and additional controls. Additionally, unlike classifier guidance, it eliminates the need to train classifiers from intermediate states to additional controls. The code is available at https://github.com/zhaoyl18/CTRL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e68c943c-eea8-4e46-8025-6b727c98b426Cited by top-tier papers4
- Training-Free Guidance Beyond Differentiability: Scalable Path Steering with Tree Search in Diffusion and Flow ModelsYingqing Guo, Yukang Yang, Hui Yuan, Mengdi WangNeurIPS 2025 · 29 citations
- Towards Better Optimization For Listwise Preference in Diffusion ModelsJiamu Bai, Xin Yu, Meilong Xu, Weitao Lu et al.ICLR 2026 · 8 citations
- Training-Free Adaptation of Diffusion Models via Doob's -TransformQijie Zhu, Zeqi Ye, Han Liu, Zhaoran Wang et al.ICML 2026 · 3 citations
- Stochastic Control for Fine-tuning Diffusion Models: Optimality, Regularity, and ConvergenceYinbin Han, Meisam Razaviyayn, Renyuan XuICML 2025
Builds on35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Offline-to-Online Reinforcement Learning with Classifier-Free Diffusion GenerationXiao Huang, Xu Liu, Enze Zhang, Tong Yu et al.ICML 2025
- Prior-Guided Diffusion Planning for Offline Reinforcement LearningDonghyeon Ki, JunHyeok Oh, Seong-Woong Shim, Byung-Jun LeeNeurIPS 2025 · 16 citations
- Training Diffusion Models Towards Diverse Image Generation with Reinforcement LearningZichen Miao, Jiang Wang, Ze Wang, Zhengyuan Yang et al.CVPR 2024 · 12 citations
- Elucidating the design space of classifier-guided diffusion generationJiajun Ma, Tianyang Hu, Wenjia Wang, Jiacheng SunICLR 2024 · 24 citations
- Understanding and Improving Training-free Loss-based Diffusion GuidanceYifei Shen, Xinyang Jiang, Yifan Yang, Yezhen Wang et al.NeurIPS 2024 · 36 citations
