Subject-driven Text-to-Image Generation via Preference-based Reinforcement Learning
Yanting Miao, William Loh, Suraj Kothawade, Pascal Poupart, Abdullah Rashwan, Yeqing Li
Abstract
Text-to-image generative models have recently attracted considerable interest, enabling the synthesis of high-quality images from textual prompts. However, these models often lack the capability to generate specific subjects from given reference images or to synthesize novel renditions under varying conditions. Methods like DreamBooth and Subject-driven Text-to-Image (SuTI) have made significant progress in this area. Yet, both approaches primarily focus on enhancing similarity to reference images and require expensive setups, often overlooking the need for efficient training and avoiding overfitting to the reference images. In this work, we present the -Harmonic reward function, which provides a reliable reward signal and enables early stopping for faster training and effective regularization. By combining the Bradley-Terry preference model, the -Harmonic reward function also provides preference labels for subject-driven generation tasks. We propose Reward Preference Optimization (RPO), which offers a simpler setup (requiring only of the negative samples used by DreamBooth) and fewer gradient steps for fine-tuning. Unlike most existing methods, our approach does not require training a text encoder or optimizing text embeddings and achieves text-image alignment by fine-tuning only the U-Net component. Empirically, -Harmonic proves to be a reliable approach for model selection in subject-driven generation tasks. Based on preference labels and early stopping validation from the -Harmonic reward function, our algorithm achieves a state-of-the-art CLIP-I score of 0.833 and a CLIP-T score of 0.314 on DreamBench.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 51c01d0a-1bd1-45b8-a78b-a0b449ab15adCited by top-tier papers6
- OmniGen2: Towards Instruction-Aligned Multimodal GenerationChenyuan Wu, Jiahao Wang, Pengfei Zheng, Ruiran Yan et al.CVPR 2026 · 231 citations
- From Broad Exploration to Stable Synthesis: Entropy-Guided Optimization for Autoregressive Image GenerationHan Song, Yucheng Zhou, Jianbing Shen, Yu ChengICLR 2026 · 9 citations
- From Competition to Synergy: Unlocking Reinforcement Learning for Subject-Driven Image GenerationZiwei Huang, Yin Shu, Hao Fang, Quanyu Long et al.ACL 2026 · 4 citations
- POCA: Pareto-Optimal Curriculum Alignment for Visual Text GenerationYaohou Fan, Qingzhong Wang, Yongsong Huang, Junyi Liu et al.CVPR 2026 · 2 citations
- GimmBO: Interactive Generative Image Model Merging via Bayesian OptimizationChenxi Liu, Selena Ling, Alec JacobsonSIGGRAPH 2026
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- Rethinking DPO-Style Diffusion Aligning FrameworksXun Wu, Shaohan Huang, Lingjie Jiang, Furu WeiICCV 2025 · 4 citations
- Margin-Aware Preference Optimization for Aligning Diffusion Models Without ReferenceJiwoo Hong, Sayak Paul, Noah Lee, Kashif Rasul et al.AAAI 2026 · 43 citations
- Scalable Ranked Preference Optimization for Text-To-Image GenerationShyamgopal Karthik, Huseyin Coskun, Zeynep Akata, Sergey Tulyakov et al.ICCV 2025 · 1 citation
- Calibrated Multi-Preference Optimization for Aligning Diffusion ModelsKyungmin Lee, Xiahong Li, Qifei Wang, Junfeng He et al.CVPR 2025
- A Dense Reward View on Aligning Text-to-Image Diffusion with PreferenceShentao Yang, Tianqi Chen, Mingyuan ZhouICML 2024 · 53 citations
