Enhancing Reward Models for High-Quality Image Generation: Beyond Text-Image Alignment
Ying Ba, Tianyu Zhang, Yalong Bai, Wenyi Mo, Tao Liang, Bing Su, Ji-Rong Wen
摘要
Contemporary image generation systems have achieved high fidelity and superior aesthetic quality beyond basic text-image alignment. However, existing evaluation frameworks have failed to evolve in parallel. This study reveals that human preference reward models fine-tuned based on CLIP and BLIP architectures have inherent flaws: they inappropriately assign low scores to images with rich details and high aesthetic value, creating a significant discrepancy with actual human aesthetic preferences. To address this issue, we design a novel evaluation score, ICT (Image-Contained-Text) score, that achieves and surpasses the objectives of text-image alignment by assessing the degree to which images represent textual content. Building upon this foundation, we further train an HP (High-Preference) score model using solely the image modality to enhance image aesthetics and detail quality while maintaining text-image alignment. Experiments demonstrate that the proposed evaluation model improves scoring accuracy by over 10% compared to existing methods, and achieves significant results in optimizing state-of-theart text-to-image models. This research provides theoretical and empirical support for evolving image generation technology toward higher-order human aesthetic preferences. Code is available at https: // github. com/ BarretBa/ ICTHP .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator OptimizationYunlong Lin, Linqing Wang, Kunjie Lin, Zixu Lin 等CVPR 2026 · 被引用 31 次
- The Image as Its Own Reward: Reinforcement Learning with Adversarial Reward for Image GenerationWeijia Mao, Hao Chen, Zhenheng Yang, Mike Zheng ShouCVPR 2026 · 被引用 10 次
- Fine-grained Image Aesthetic Assessment: Learning Discriminative Scores from Relative RanksZhichao Yang, Jianjie Wang, Zhixianhe Zhang, Pangu Xie 等CVPR 2026 · 被引用 5 次
- PreferThinker: Reasoning-based Personalized Image Preference AssessmentShengqi Xu, Xinpeng Zhou, Yabo Zhang, Ming Liu 等ICLR 2026
- Pareto-Guided Optimal Transport for Multi-Reward AlignmentYing Ba, Tianyu Zhang, Mohan Zhou, Yalong Bai 等ICML 2026
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
相关 Paper
- Learning Multi-Dimensional Human Preference for Text-to-Image GenerationSixian Zhang, Bohan Wang, Junqiang Wu, Yan Li 等CVPR 2024 · 被引用 15 次
- LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis EvaluationYujie Lu, Xianjun Yang, Xiujun Li, Xin Eric Wang 等NeurIPS 2023 · 被引用 119 次
- Aligning Multi-Character Narrative Image Generation with Multi-Aspect Human PreferencesZiyi Gao, Zhipeng Wei, Jingjing Chen, Zhiyu Tan 等CVPR 2026
- EvalMuse-40K: A Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Alignment EvaluationShuhao Han, Haotian Fan, Jiachen Fu, Liang Li 等AAAI 2026 · 被引用 1 次
- Science-T2I: Addressing Scientific Illusions in Image SynthesisJialuo Li, Wenhao Chai, Xingyu Fu, Haiyang Xu 等CVPR 2025
