TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering
Yushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang, Mari Ostendorf, Ranjay Krishna, Noah A. Smith
Abstract
Despite thousands of researchers, engineers, and artists actively working on improving text-to-image generation models, systems often fail to produce images that accurately align with the text inputs. We introduce TIFA (Text-to-Image Faithfulness evaluation with question Answering), an automatic evaluation metric that measures the faithfulness of a generated image to its text input via visual question answering (VQA). Specifically, given a text input, we automatically generate several question-answer pairs using a language model. We calculate image faithfulness by checking whether existing VQA models can answer these questions using the generated image. TIFA is a reference-free metric that allows for fine-grained and interpretable evaluations of generated images. TIFA also has better correlations with human judgments than existing metrics. Based on this approach, we introduce TIFA v1.0, a benchmark consisting of 4K diverse text inputs and 25K questions across 12 categories (object, counting, etc.). We present a comprehensive evaluation of existing text-to-image models using TIFA v1.0 and highlight the limitations and challenges of current models. For instance, we find that current text-to-image models, despite doing well on color and material, still struggle in counting, spatial relations, and composing multiple objects. We hope our benchmark will help carefully measure the research progress in text-to-image synthesis and provide valuable insights for further research. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d38b879c-86d3-40b4-a27f-bb01bbeae630Cited by top-tier papers142
- What You See is What You Read? Improving Text-Image Alignment EvaluationMichal Yarom, Yonatan Bitton, Soravit Changpinyo, Roee Aharoni et al.NeurIPS 2023 · 147 citations
- Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image GenerationJaemin Cho, Yushi Hu, Jason M. Baldridge, Roopal Garg et al.ICLR 2024 · 139 citations
- Navigating Text-To-Image Customization: From LyCORIS Fine-Tuning to Model EvaluationShih-Ying Yeh, Yu-Guan Hsieh, Zhidong Gao, Bernard B. W. Yang et al.ICLR 2024 · 133 citations
- Vision-by-Language for Training-Free Compositional Image RetrievalShyamgopal Karthik, Karsten Roth, Massimiliano Mancini, Zeynep AkataICLR 2024 · 120 citations
- The Generative AI Paradox: "What It Can Create, It May Not Understand"Peter West, Ximing Lu, Nouha Dziri, Faeze Brahman et al.ICLR 2024 · 116 citations
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Who Evaluates the Evaluations? Objectively Scoring Text-to-Image Prompt Coherence Metrics with T2IScoreScore (TS2)Michael Saxon, Fatima Jahara, Mahsa Khoshnoodi, Yujie Lu et al.NeurIPS 2024 · 19 citations
- Evaluating Image Hallucination in Text-to-Image Generation with Question-AnsweringYoungsun Lim, Hojun Choi, Hyunjung ShimAAAI 2025 · 16 citations
- Revisiting text-to-image evaluation with Gecko: on metrics, prompts, and human ratingOlivia Wiles, Chuhan Zhang, Isabela Albuquerque, Ivana Kajic et al.ICLR 2025
- Toward Verifiable and Reproducible Human Evaluation for Text-to-Image GenerationMayu Otani, Riku Togashi, Yu Sawai, Ryosuke Ishigami et al.CVPR 2023
- Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchySimon Ging, María Alejandra Bravo, Thomas BroxICLR 2024 · 24 citations
