ACL2026
REVEALER: Reinforcement-Guided Visual Reasoning for Element-Level Text-Image Alignment Evaluation
Fulin Shi, Wenyi Xiao, Bin Chen, Liang Ding, Leilei Gan
被引用 3 次
摘要
Evaluating the alignment between textual prompts and generated images is critical for ensuring the reliability and usability of textto-image (T2I) models. However, most existing evaluation methods rely on coarsegrained metrics or static Question Answering (QA) pipelines, which lack fine-grained interpretability and struggle to reflect human preferences. To address this, we propose REVEALER, a reinforcement-guided visual reasoning framework for element-level textto-image alignment evaluation. Adopting a structured "grounding-reasoning-conclusion" paradigm, our method enables Multimodal Large Language Models (MLLMs) to explicitly localize semantic elements and derive interpretable alignment judgments. We optimize the model via Group Relative Policy Optimization (GRPO) using a multi-dimensional reward function that targets format compliance, localization precision, and alignment accuracy. Extensive experiments confirm that REVEALER achieves state-of-the-art results across four benchmarks. Notably, on EvalMuse-40K, it surpasses the strong proprietary Gemini 3 Pro and Training-based baselines with absolute accuracy gains of +4.0% and +13.1%, respectively. Ablation studies further demonstrate the efficacy of our method, contributing a cumulative 19.4%