Curriculum Direct Preference Optimization for Diffusion and Consistency Models
Florinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu, Nicu Sebe, Mubarak Shah
Abstract
Direct Preference Optimization (DPO) has been proposed as an effective and efficient alternative to reinforcement learning from human feedback (RLHF). In this paper, we propose a novel and enhanced version of DPO based on curriculum learning for text-to-image generation. Our method is divided into two training stages. First, a ranking of the examples generated for each prompt is obtained by employing a reward model. Then, increasingly difficult pairs of examples are sampled and provided to a text-to-image generative (diffusion or consistency) model. Generated samples that are far apart in the ranking are considered to form easy pairs, while those that are close in the ranking form hard pairs. In other words, we use the rank difference between samples as a measure of difficulty. The sampled pairs are split into batches according to their difficulty levels, which are gradually used to train the generative model. Our approach, Curriculum DPO, is compared against state-of-the-art fine-tuning approaches on nine benchmarks, outperforming the competing methods in terms of text alignment, aesthetics and human preference. Our code is available at https://github. com/CroitoruAlin/Curriculum-DPO .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 32a0889a-c8db-476a-8723-374517b2e92dCited by top-tier papers11
- Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image ModelsZengbin Wang, Xuecai Hu, Yong Wang, Feng Xiong et al.ICLR 2026 · 13 citations
- Dual-Difficulty Curriculum Learning for Direct Preference OptimizationMengyang Li, Haozhan Geng, Zhong Zhang, Shuang LiuKDD 2026 · 5 citations
- POCA: Pareto-Optimal Curriculum Alignment for Visual Text GenerationYaohou Fan, Qingzhong Wang, Yongsong Huang, Junyi Liu et al.CVPR 2026 · 2 citations
- Offline Preference Optimization for Rectified Flow with Noise-Tracked PairsYunhong Lu, Qichao Wang, Hengyuan Cao, Xiaoyin Xu et al.ICML 2026 · 1 citation
- Test-Time Preference Optimization for Image RestorationBingchen Li, Xin Li, Jiaqi Xu, Jiaming Guo et al.AAAI 2026 · 1 citation
Builds on34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
Related papers
- Diffusion Model Alignment Using Direct Preference OptimizationBram Wallace, Meihua Dang, Rafael Rafailov, Linqi Zhou et al.CVPR 2024 · 89 citations
- DSPO: Direct Score Preference Optimization for Diffusion Model AlignmentHuaisheng Zhu, Teng Xiao, Vasant G. HonavarICLR 2025
- Rethinking DPO-Style Diffusion Aligning FrameworksXun Wu, Shaohan Huang, Lingjie Jiang, Furu WeiICCV 2025 · 4 citations
- Ranking-based Preference Optimization for Diffusion Models from Implicit User FeedbackYi-Lun Wu, Bo-Kai Ruan, Chiang Tseng, Hong-Han ShuaiNeurIPS 2025 · 3 citations
- InPO: Inversion Preference Optimization with Reparametrized DDIM for Efficient Diffusion Model AlignmentYunhong Lu, Qichao Wang, Hengyuan Cao, Xierui Wang et al.CVPR 2025
