Calibrated Multi-Preference Optimization for Aligning Diffusion Models
Kyungmin Lee, Xiahong Li, Qifei Wang, Junfeng He, Junjie Ke, Ming-Hsuan Yang, Irfan Essa, Jinwoo Shin, Feng Yang, Yinxiao Li
Abstract
Aligning text-to-image (T2I) diffusion models with preference optimization is valuable for human-annotated datasets, but the heavy cost of manual data collection limits scalability. Using reward models offers an alternative, however, current preference optimization methods fall short in exploiting the rich information, as they only consider pairwise preference distribution. Furthermore, they lack generalization to multi-preference scenarios and struggle to handle inconsistencies between rewards. To address this, we present Calibrated Preference Optimization (CaPO), a novel method to align T2I diffusion models by incorporating the general preference from multiple reward models without human annotated data. The core of our approach involves a reward calibration method to approximate the general preference by computing the expected win-rate against the samples generated by the pretrained models. Additionally, we propose a frontier-based pair selection method that effectively manages the multi-preference distribution by selecting pairs from Pareto frontiers. Finally, we use regression loss to fine-tune diffusion models to match the difference between calibrated rewards of a selected pair. Experimental results show that CaPO consistently outperforms prior methods, such as Direct Preference Optimization (DPO), in both single and multi-reward settings validated by evaluation on T2I benchmarks, including GenEval and T2I-Compbench.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 463235e4-e588-4004-b393-74007b64d9e1Cited by top-tier papers23
- Aligning Text to Image in Diffusion Models is Easier Than You ThinkJaa-Yeon Lee, Byunghee Cha, Jeongsol Kim, Jong Chul YeNeurIPS 2025 · 23 citations
- Follow-Your-Preference: Towards Preference-Aligned Image InpaintingYutao Shen, Junkun Yuan, Toru Aonishi, Hideki Nakayama et al.ICLR 2026 · 21 citations
- Decoupled MeanFlow: Turning Flow Models into Flow Maps for Accelerated SamplingKyungmin Lee, Sihyun Yu, Jinwoo ShinICLR 2026 · 18 citations
- Flash-DMD: Towards High-Fidelity Few-Step Image Generation with Efficient Distillation and Joint Reinforcement LearningGuanjie Chen, Shirui Huang, Yifu Sun, Kai Liu et al.CVPR 2026 · 17 citations
- Multi-Objective Preference Optimization: Improving Human Alignment of Generative ModelsAkhil Agnihotri, Rahul Jain, Deepak Ramachandran, Zheng WenICML 2026 · 15 citations
Builds on36
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- Aligning Diffusion Models by Optimizing Human UtilityShufan Li, Konstantinos Kallidromitis, Akash Gokul, Yusuke Kato et al.NeurIPS 2024 · 117 citations
- Rethinking DPO-Style Diffusion Aligning FrameworksXun Wu, Shaohan Huang, Lingjie Jiang, Furu WeiICCV 2025 · 4 citations
- Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human PreferencesYunhong Lu, Qichao Wang, Hengyuan Cao, Xiaoyin Xu et al.ICML 2025
- Scalable Ranked Preference Optimization for Text-To-Image GenerationShyamgopal Karthik, Huseyin Coskun, Zeynep Akata, Sergey Tulyakov et al.ICCV 2025 · 1 citation
- Curriculum Direct Preference Optimization for Diffusion and Consistency ModelsFlorinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu, Nicu Sebe et al.CVPR 2025
