e-ViL: A Dataset and Benchmark for Natural Language Explanations in Vision-Language Tasks
Maxime Kayser, Oana-Maria Camburu, Leonard Salewski, Cornelius Emde, Virginie Do, Zeynep Akata, Thomas Lukasiewicz
Abstract
Recently, there has been an increasing number of efforts to introduce models capable of generating natural language explanations (NLEs) for their predictions on vision-language (VL) tasks. Such models are appealing, because they can provide human-friendly and comprehensive explanations. However, there is a lack of comparison between existing methods, which is due to a lack of re-usable evaluation frameworks and a scarcity of datasets. In this work, we introduce e-ViL and e-SNLI-VE. e-ViL is a benchmark for explainable vision-language tasks that establishes a unified evaluation framework and provides the first comprehensive comparison of existing approaches that generate NLEs for VL tasks. It spans four models and three datasets and both automatic metrics and human evaluation are used to assess modelgenerated explanations. e-SNLI-VE is currently the largest existing VL dataset with NLEs (over 430k instances). We also propose a new model that combines UNITER [15] , which learns joint embeddings of images and text, and GPT-2 [38], a pre-trained language model that is well-suited for text generation. It surpasses the previous state of the art by a large margin across all datasets. Code and data are available here: https://github.com/maximek3/e-ViL . Hypothesis: A man and woman inside a church. Textual premise: A man and woman getting married. Original label: Neutral Caption #2: A man and woman that is holding flowers smile in the sunlight. Caption #4: A happy couple enjoying their open air wedding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers14
- Large Language Models are Visual Reasoning CoordinatorsLiangyu Chen, Bo Li, Sheng Shen, Jingkang Yang et al.NeurIPS 2023 · 108 citations
- NLX-GPT: A Model for Natural Language Explanations in Vision and Vision-Language TasksFawaz Sammani, Tanmoy Mukherjee, Nikos DeligiannisCVPR 2022 · 46 citations
- What Factors Affect Multi-Modal In-Context Learning? An In-Depth ExplorationLibo Qin, Qiguang Chen, Hao Fei, Zhi Chen et al.NeurIPS 2024 · 37 citations
- Symbolic Chain-of-Thought Distillation: Small Models Can Also "Think" Step-by-StepLiunian Harold Li, Jack Hessel, Youngjae Yu, Xiang Ren et al.ACL 2023 · 34 citations
- Beyond task performance: evaluating and reducing the flaws of large multimodal models with in-context-learningMustafa Shukor, Alexandre Ramé, Corentin Dancette, Matthieu CordICLR 2024 · 31 citations
Builds on6
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Interpreting Interpretability: Understanding Data Scientists' Use of Interpretability Tools for Machine LearningHarmanpreet Kaur, Harsha Nori, Samuel Jenkins, Rich Caruana et al.CHI 2020 · 541 citations
- Abductive Commonsense ReasoningChandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi et al.ICLR 2020 · 521 citations
- Generating Fact Checking ExplanationsPepa Atanasova, Jakob Grue Simonsen, Christina Lioma, Isabelle AugensteinACL 2020 · 130 citations
- NILE : Natural Language Inference with Faithful Natural Language ExplanationsSawan Kumar, Partha P. TalukdarACL 2020 · 15 citations
Related papers
- V2P-Bench: Evaluating Video-Language Understanding with Visual Prompts for Better Human-Model InteractionYiming Zhao, Yu Zeng, Yukun Qi, YaoYang Liu et al.ICLR 2026 · 8 citations
- Unsupervised Explanation Generation via Correct InstantiationsSijie Cheng, Zhiyong Wu, Jiangjie Chen, Zhixing Li et al.AAAI 2023 · 6 citations
- ELITE: Enhanced Language-Image Toxicity Evaluation for SafetyWonjun Lee, Doehyeon Lee, Eugene Choi, Sangyoon Yu et al.ICML 2025
- CLEVR-Implicit: A Diagnostic Dataset for Implicit Reasoning in Referring Expression ComprehensionJingwei Zhang, Xin Wu, Yi CaiEMNLP 2023 · 1 citation
- ReForm-Eval: Evaluating Large Vision Language Models via Unified Re-Formulation of Task-Oriented BenchmarksZejun Li, Ye Wang, Mengfei Du, Qingwen Liu et al.ACM MM 2024 · 1 citation
