FIND: Fine-tuning Initial Noise Distribution with Policy Optimization for Diffusion Models
Changgu Chen, Libing Yang, Xiaoyan Yang, Lianggangxu Chen, Gaoqi He, Changbo Wang, Yang Li
摘要
In recent years, large-scale pre-trained diffusion models have demonstrated their outstanding capabilities in image and video generation tasks. However, existing models tend to produce visual objects commonly found in the training dataset, which diverges from user input prompts. The underlying reason behind the inaccurate generated results lies in the model's difficulty in sampling from specific intervals of the initial noise distribution corresponding to the prompt. Moreover, it is challenging to directly optimize the initial distribution, given that the diffusion process involves multiple denoising steps. In this paper, we introduce a Fine-tuning Initial Noise Distribution (FIND) framework with policy optimization, which unleashes the powerful potential of pre-trained diffusion networks by directly optimizing the initial distribution to align the generated contents with user-input prompts. To this end, we first reformulate the diffusion denoising procedure as a one-step Markov decision process and employ policy optimization to directly optimize the initial distribution. In addition, a dynamic reward calibration module is proposed to ensure training stability during optimization. Furthermore, we introduce a ratio clipping algorithm to utilize historical data for network training and prevent the optimized distribution from deviating too far from the original policy to restrain excessive optimization magnitudes. Extensive experiments demonstrate the effectiveness of our method in both text-to-image and text-to-video tasks, surpassing SOTA methods in achieving consistency between prompts and the generated content. Our method achieves 10 times faster than the SOTA approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Physics-Informed Distillation of Diffusion Models for PDE-Constrained GenerationYi Zhang, Peng Wang, Difan ZouICML 2026 · 被引用 9 次
- Adjusting Initial Noise to Mitigate Memorization in Text-to-Image Diffusion ModelsHyeonggeun Han, Sehwan Kim, Hyungjun Joo, Sangwoo Hong 等NeurIPS 2025 · 被引用 7 次
- Noise Matters: Optimizing Matching Noise for Diffusion ClassifiersYanghao Wang, Long ChenNeurIPS 2025 · 被引用 7 次
- FeRA: Frequency-Energy Constrained Routing for Effective Diffusion Adaptation Fine-TuningBo Yin, Xiaobin Hu, Xingyu Zhou, Yu HE 等ICML 2026 · 被引用 5 次
- Antithetic Noise in Diffusion ModelsJing Jia, Sifan Liu, Bowen Song, Wei Yuan 等ICLR 2026 · 被引用 4 次
它引用的顶会 Paper32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
相关 Paper
- Using Human Feedback to Fine-tune Diffusion Models without Any Reward ModelKai Yang, Jian Tao, Jiafei Lyu, Chunjiang Ge 等CVPR 2024 · 被引用 34 次
- Score as Action: Fine Tuning Diffusion Generative Models by Continuous-time Reinforcement LearningHanyang Zhao, Haoxian Chen, Ji Zhang, David D. Yao 等ICML 2025
- Prompt-A-Video: Prompt your Video Diffusion Model via Preference-Aligned LLMYatai Ji, Jiacheng Zhang, Jie Wu, Shilong Zhang 等ICCV 2025 · 被引用 3 次
- Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion modelsKyungmin Lee, Sangkyung Kwak, Kihyuk Sohn, Jinwoo ShinNeurIPS 2024 · 被引用 13 次
- Inference-Time Alignment of Diffusion Models with Direct Noise OptimizationZhiwei Tang, Jiangweizhi Peng, Jiasheng Tang, Mingyi Hong 等ICML 2025
