Enhancing Reward Models for High-Quality Image Generation: Beyond Text-Image Alignment
Ying Ba, Tianyu Zhang, Yalong Bai, Wenyi Mo, Tao Liang, Bing Su, Ji-Rong Wen
Abstract
Contemporary image generation systems have achieved high fidelity and superior aesthetic quality beyond basic text-image alignment. However, existing evaluation frameworks have failed to evolve in parallel. This study reveals that human preference reward models fine-tuned based on CLIP and BLIP architectures have inherent flaws: they inappropriately assign low scores to images with rich details and high aesthetic value, creating a significant discrepancy with actual human aesthetic preferences. To address this issue, we design a novel evaluation score, ICT (Image-Contained-Text) score, that achieves and surpasses the objectives of text-image alignment by assessing the degree to which images represent textual content. Building upon this foundation, we further train an HP (High-Preference) score model using solely the image modality to enhance image aesthetics and detail quality while maintaining text-image alignment. Experiments demonstrate that the proposed evaluation model improves scoring accuracy by over 10% compared to existing methods, and achieves significant results in optimizing state-of-theart text-to-image models. This research provides theoretical and empirical support for evolving image generation technology toward higher-order human aesthetic preferences. Code is available at https: // github. com/ BarretBa/ ICTHP .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f426a321-18d5-4135-bc18-20ae90c51dc3Cited by top-tier papers5
- JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator OptimizationYunlong Lin, Linqing Wang, Kunjie Lin, Zixu Lin et al.CVPR 2026 · 31 citations
- The Image as Its Own Reward: Reinforcement Learning with Adversarial Reward for Image GenerationWeijia Mao, Hao Chen, Zhenheng Yang, Mike Zheng ShouCVPR 2026 · 10 citations
- Fine-grained Image Aesthetic Assessment: Learning Discriminative Scores from Relative RanksZhichao Yang, Jianjie Wang, Zhixianhe Zhang, Pangu Xie et al.CVPR 2026 · 5 citations
- PreferThinker: Reasoning-based Personalized Image Preference AssessmentShengqi Xu, Xinpeng Zhou, Yabo Zhang, Ming Liu et al.ICLR 2026
- Pareto-Guided Optimal Transport for Multi-Reward AlignmentYing Ba, Tianyu Zhang, Mohan Zhou, Yalong Bai et al.ICML 2026
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Learning Multi-Dimensional Human Preference for Text-to-Image GenerationSixian Zhang, Bohan Wang, Junqiang Wu, Yan Li et al.CVPR 2024 · 15 citations
- LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis EvaluationYujie Lu, Xianjun Yang, Xiujun Li, Xin Eric Wang et al.NeurIPS 2023 · 119 citations
- Aligning Multi-Character Narrative Image Generation with Multi-Aspect Human PreferencesZiyi Gao, Zhipeng Wei, Jingjing Chen, Zhiyu Tan et al.CVPR 2026
- EvalMuse-40K: A Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Alignment EvaluationShuhao Han, Haotian Fan, Jiachen Fu, Liang Li et al.AAAI 2026 · 1 citation
- Science-T2I: Addressing Scientific Illusions in Image SynthesisJialuo Li, Wenhao Chai, Xingyu Fu, Haiyang Xu et al.CVPR 2025
