What You See is What You Read? Improving Text-Image Alignment Evaluation
Michal Yarom, Yonatan Bitton, Soravit Changpinyo, Roee Aharoni, Jonathan Herzig, Oran Lang, Eran Ofek, Idan Szpektor
摘要
Automatically determining whether a text and a corresponding image are semantically aligned is a significant challenge for vision-language models, with applications in generative text-to-image and image-to-text tasks. In this work, we study methods for automatic text-image alignment evaluation. We first introduce SeeTRUE: a comprehensive evaluation set, spanning multiple datasets from both text-to-image and image-to-text generation tasks, with human judgements for whether a given text-image pair is semantically aligned. We then describe two automatic methods to determine alignment: the first involving a pipeline based on question generation and visual question answering models, and the second employing an end-to-end classification approach by finetuning multimodal pretrained models. Both methods surpass prior approaches in various text-image alignment tasks, with significant improvements in challenging cases that involve complex composition or unnatural images. Finally, we demonstrate how our approaches can localize specific misalignments between an image and a given text, and how they can be used to automatically re-rank candidates in text-to-image generation. 1 * Equal contribution. Yonatan participated in this work as part of an internship at Google Research. 1 Data and code are attached to this submission. Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper46
- MMICL: Empowering Vision-language Model with Multi-Modal In-Context LearningHaozhe Zhao, Zefan Cai, Shuzheng Si, Xiaojian Ma 等ICLR 2024 · 被引用 206 次
- VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video GenerationHritik Bansal, Clark Peng, Yonatan Bitton, Roman Goldenberg 等ICLR 2026 · 被引用 146 次
- Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image GenerationJaemin Cho, Yushi Hu, Jason M. Baldridge, Roopal Garg 等ICLR 2024 · 被引用 139 次
- VideoPrism: A Foundational Visual Encoder for Video UnderstandingLong Zhao, Nitesh Bharadwaj Gundavarapu, Liangzhe Yuan, Hao Zhou 等ICML 2024 · 被引用 91 次
- Visual Programming for Step-by-Step Text-to-Image Generation and EvaluationJaemin Cho, Abhay Zala, Mohit BansalNeurIPS 2023 · 被引用 62 次
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- Understanding ME? Multimodal Evaluation for Fine-grained Visual CommonsenseZhecan Wang, Haoxuan You, Yicheng He, Wenhao Li 等EMNLP 2022 · 被引用 2 次
- TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question AnsweringYushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang 等ICCV 2023 · 被引用 400 次
- Divide, Evaluate, and Refine: Evaluating and Improving Text-to-Image Alignment with Iterative VQA FeedbackJaskirat Singh, Liang ZhengNeurIPS 2023 · 被引用 48 次
- Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive EvaluationHong-Tao Yu, Yuxin Peng, Serge J. Belongie, Xiu-Shen WeiICLR 2026 · 被引用 21 次
- Discriminative Probing and Tuning for Text-to-Image GenerationLeigang Qu, Wenjie Wang, Yongqi Li, Hanwang Zhang 等CVPR 2024 · 被引用 7 次
