Diffusion Negative Preference Optimization Made Simple
Joshua Tian Jin Tee, Hee Suk Yoon, Sunjae Yoon, Tri Ton, Chang Yoo
摘要
Classifier-Free Guidance (CFG) improves diffusion sampling by encouraging conditional generations while discouraging unconditional ones. Existing preference alignment methods, however, focus only on positive preference pairs, limiting their ability to actively suppress undesirable outputs. Diffusion Negative Preference Optimization (Diff-NPO) approaches this limitation by introducing a separate negative model trained with inverted labels, allowing it to capture signals for suppressing undesirable generations. However, this design comes with two key drawbacks. First, maintaining two distinct models throughout training and inference substantially increases computational cost, making the approach less practical. Second, at inference time, Diff-NPO relies on weight merging between the positive and negative models, a process that dilutes the learned negative alignment and undermines its effectiveness. To overcome these issues, we introduce Diff-SNPO, a single-network framework that jointly learns from both positive and negative preferences. Our method employs a bounded preference objective to prevent winner-likelihood collapse, ensuring stable optimization. Diff-SNPO delivers strong alignment performance with significantly lower computational overhead, showing that explicit negative preference modeling can be simple, stable, and efficient within a unified diffusion framework. Code and models are available at https://github.com/JoshuaTTJ/DiffSNPO..
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
相关 Paper
- Diffusion-NPO: Negative Preference Optimization for Better Preference Aligned Generation of Diffusion ModelsFu-Yun Wang, Yunhao Shui, Jingtan Piao, Keqiang Sun 等ICLR 2025
- Self-NPO: Data-Free Diffusion Model Enhancement via Truncated Diffusion Fine-TuningFu-Yun Wang, Keqiang Sun, Yao Teng, Xihui Liu 等AAAI 2026 · 被引用 1 次
- ContrastiveCFG: Guiding Diffusion Sampling by Contrasting Positive and Negative ConceptsJinho Chang, Changsun Lee, Hyungjin Chung, Jong Chul YEICML 2026
- Normalized Attention Guidance: Universal Negative Guidance for Diffusion ModelsDar-Yen Chen, Hmrishav Bandyopadhyay, Kai Zou, Yi-Zhe SongNeurIPS 2025 · 被引用 17 次
- Inversion-DPO: Precise and Efficient Post-Training for Diffusion ModelsZejian Li, Yize Li, Chenye Meng, Zhongni Liu 等ACM MM 2025
