Entropy-Adaptive Diffusion Policy Optimization with Dynamic Step Alignment
Renye Yan, Jikang Cheng, Yaozhong Gan, Shikun Sun, You Wu, Yunfan Yang, Ling Liang, Jinlong Lin, Yeshuang Zhu, Jie Zhou, Jinchao Zhang, Junliang Xing
Abstract
While fine-tuning diffusion models with reinforcement learning (RL) has demonstrated effectiveness in directly optimizing downstream objectives, existing RL frameworks are prone to overfitting the rewards, leading to outputs that deviate from the true data distribution and exhibit reduced diversity. To address this issue, we introduce entropy as a quantitative measure to enhance the exploratory capacity of diffusion models' denoising policies. We propose an adaptive mechanism that dynamically adjusts the application and magnitude of entropy and regularization, guided by real-time quality estimation of intermediate noised states. Theoretically, we prove the convergence of our entropyenhanced policy optimization and establish two critical properties: 1) global entropy increases through training, ensuring robust exploration capabilities, and 2) entropy systematically decreases during the denoising process, enabling a phase transition from early-stage diversity promotion to late-stage distributional fidelity. Building on this foundation, we propose a plug-and-play RL module that adaptively controls entropy and optimizes denoising steps. Extensive evaluations demonstrate theoretical soundness and empirical robustness of our method, achieving state-ofthe-art quality-diversity trade-offs across benchmarks. Notably, our framework significantly improves the rewards and reduces denoising steps in training by up to 40%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5b6f1a15-5e16-42cf-bf40-d2cd81a7cdeeCited by top-tier papers3
- Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models?Renye Yan, Jikang Cheng, Shikun Sun, Yi Sun et al.CVPR 2026 · 9 citations
- TEAM: Temporal–Spatial Consistency Guided Expert Activation for MoE Diffusion Language Model AccelerationLINYE WEI, Zixiang Luo, Pingzhi Tang, Meng LiICML 2026 · 7 citations
- FastSESR: Fast Scene-level Explicit Surface ReconstructionJueqi Liu, Xuechao Zou, Congyan LangICML 2026
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
Related papers
- Training Diffusion Models Towards Diverse Image Generation with Reinforcement LearningZichen Miao, Jiang Wang, Ze Wang, Zhengyuan Yang et al.CVPR 2024 · 12 citations
- Diffusion Fine-Tuning via Reparameterized Policy Gradient of the Soft Q-FunctionHyeongyu Kang, Jaewoo Lee, Woocheol Shin, Kiyoung Om et al.ICLR 2026 · 5 citations
- Diffusion Controller: Framework, Algorithms and ParameterizationTong Yang, Moonkyung Ryu, Chih-wei Hsu, Guy Tennenholtz et al.ICML 2026
- Entropy-regularized Diffusion Policy with Q-Ensembles for Offline Reinforcement LearningRuoqi Zhang, Ziwei Luo, Jens Sjölund, Thomas B. Schön et al.NeurIPS 2024 · 43 citations
- Training Diffusion Models with Reinforcement LearningKevin Black, Michael Janner, Yilun Du, Ilya Kostrikov et al.ICLR 2024 · 816 citations
