REVEALER: Reinforcement-Guided Visual Reasoning for Element-Level Text-Image Alignment Evaluation
Fulin Shi, Wenyi Xiao, Bin Chen, Liang Ding, Leilei Gan
Abstract
Evaluating the alignment between textual prompts and generated images is critical for ensuring the reliability and usability of textto-image (T2I) models. However, most existing evaluation methods rely on coarsegrained metrics or static Question Answering (QA) pipelines, which lack fine-grained interpretability and struggle to reflect human preferences. To address this, we propose REVEALER, a reinforcement-guided visual reasoning framework for element-level textto-image alignment evaluation. Adopting a structured "grounding-reasoning-conclusion" paradigm, our method enables Multimodal Large Language Models (MLLMs) to explicitly localize semantic elements and derive interpretable alignment judgments. We optimize the model via Group Relative Policy Optimization (GRPO) using a multi-dimensional reward function that targets format compliance, localization precision, and alignment accuracy. Extensive experiments confirm that REVEALER achieves state-of-the-art results across four benchmarks. Notably, on EvalMuse-40K, it surpasses the strong proprietary Gemini 3 Pro and Training-based baselines with absolute accuracy gains of +4.0% and +13.1%, respectively. Ablation studies further demonstrate the efficacy of our method, contributing a cumulative 19.4%
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c40070d0-436e-49f3-992f-7f5e3b60028bCited by top-tier papers1
Ask how each one uses itBuilds on21
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- EvalMuse-40K: A Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Alignment EvaluationShuhao Han, Haotian Fan, Jiachen Fu, Liang Li et al.AAAI 2026 · 1 citation
- RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement LearningMingrui Wu, Lu Wang, Pu Zhao, Fangkai Yang et al.ICLR 2026 · 19 citations
- GenAlign: Towards Unified Alignment Framework of MLLMs via Generative Reward ModelJingyu Zhang, Kun Yang, Ming Wen, jiawei zhao et al.ICML 2026
- SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy OptimizationXuankun Rong, Wenke Huang, Tingfeng Wang, Daiguo Zhou et al.CVPR 2026 · 13 citations
- LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis EvaluationYujie Lu, Xianjun Yang, Xiujun Li, Xin Eric Wang et al.NeurIPS 2023 · 119 citations
