Confidence-aware Reward Optimization for Fine-tuning Text-to-Image Models
Kyuyoung Kim, Jongheon Jeong, Minyong An, Mohammad Ghavamzadeh, Krishnamurthy Dj Dvijotham, Jinwoo Shin, Kimin Lee
Abstract
Fine-tuning text-to-image models with reward functions trained on human feedback data has proven effective for aligning model behavior with human intent. However, excessive optimization with such reward models, which serve as mere proxy objectives, can compromise the performance of fine-tuned models, a phenomenon known as reward overoptimization. To investigate this issue in depth, we introduce the Text-Image Alignment Assessment (TIA2) benchmark, which comprises a diverse collection of text prompts, images, and human annotations. Our evaluation of several state-of-the-art reward models on this benchmark reveals their frequent misalignment with human assessment. We empirically demonstrate that overoptimization occurs notably when a poorly aligned reward model is used as the fine-tuning objective. To address this, we propose TextNorm, a simple method that enhances alignment based on a measure of reward model confidence estimated across a set of semantically contrastive text prompts. We demonstrate that incorporating the confidence-calibrated rewards in fine-tuning effectively reduces overoptimization, resulting in twice as many wins in human evaluation for text-image alignment compared against the baseline reward models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMsRui Yang, Ruomeng Ding, Yong Lin, Huan Zhang et al.NeurIPS 2024 · 157 citations
- DiffExp: Efficient Exploration in Reward Fine-tuning for Text-to-Image Diffusion ModelsDaewon Chae, June Suk Choi, Jinkyu Kim, Kimin LeeAAAI 2025 · 3 citations
- Bias at the End of the ScoreSalma Abdel Magid, Grace Guo, Esin Tureci, Amaya Dharmasiri et al.CVPR 2026 · 1 citation
- Enhancing Ligand Validity and Affinity in Structure-Based Drug Design with Multi-Reward OptimizationSeungbeom Lee, Munsun Jo, Jungseul Ok, Dongwoo KimICML 2025
- Scaling Inference Time Compute for Diffusion ModelsNanye Ma, Shangyuan Tong, Haolin Jia, Hexiang Hu et al.CVPR 2025
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
Related papers
- Calibrated Multi-Preference Optimization for Aligning Diffusion ModelsKyungmin Lee, Xiahong Li, Qifei Wang, Junfeng He et al.CVPR 2025
- Aligning Diffusion Models by Optimizing Human UtilityShufan Li, Konstantinos Kallidromitis, Akash Gokul, Yusuke Kato et al.NeurIPS 2024 · 117 citations
- PromptEnhancer: Taming Your Rewriter for Text-to-Image Generation via Fine-Grained RewardLinqing Wang, Zhiyong Xu, Ximing Xing, Yiji Cheng et al.CVPR 2026
- Style-GRPO: Semantic-Aware Preference Optimization for Image Style Transfer Guided by Reward ModelingJianbin Zhao, Chaoran Feng, Miao Yu, Yingtao Li et al.CVPR 2026
- RubricBench: Aligning Model-Generated Rubrics with Human StandardsJunyi Zhou, Qiyuan Zhang, Yufei Wang, Fuyuan Lyu et al.ACL 2026 · 7 citations
