SCIEval: Evaluating and Benchmarking the Faithfulness of Scientific Image Generation and Interpretation with Large Multimodal Models
Guanghui Ye, Huan Zhao, Zhixue Zhao, Tengfei Ma, Kehan Wang, Steffen Eger, Zhihua Jiang
Abstract
Scientific images often require accurate numerical representations and correct object attributes. However, current faithfulness metrics are primarily tailored toward photorealistic, real-life imagery, rendering them ill-suited for scientific image evaluation. To address this gap, we introduce a novel evaluation model, SCIEval (SCientific Image Evaluation), which aims to capture faithfulness through three key dimensions: (i) Relevance, measuring overall textimage correspondence; (ii) Accuracy, examining the technical details of scientific objects; and (iii) Explainability, which isolates unfaithful elements within the generated content. To address these dimensions, we curate a specialized dataset of scientific text-image pairs to train three evaluation modules. For the Relevance and Accuracy modules, we propose a CLIP-based strategy that enhances scientific image perception through intra-and cross-modal contrastive learning. Concurrently, the Explainability module is developed by fine-tuning a high-performance Large Multimodal Model (LMM) using supervised rationale signals. Finally, we present SCIEval-Bench, a human-annotated evaluation benchmark consisting of 3,000 samples for scientific textto-image and 3,000 samples for scientific image captioning. Extensive experiments on SCIEval-Bench demonstrate that our SCIEval model is significantly more reliable than 24 competing models-including GPT-4o-exhibiting a superior correlation with human judgments. Project page: https://SCIEval.github.io
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e5d5d77-3997-4012-bd32-f465745f42e3Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
Related papers
- ScImage: How good are multimodal large language models at scientific text-to-image generation?Leixin Zhang, Steffen Eger, Yinjie Cheng, Weihe Zhai et al.ICLR 2025
- SciEval: A Multi-Level Large Language Model Evaluation Benchmark for Scientific ResearchLiangtai Sun, Yang Han, Zihan Zhao, Da Ma et al.AAAI 2024 · 150 citations
- Automatic and Reliable Evaluation for Academic Caption-to-Figure Generation with LMMsGuanghui Ye, Huan Zhao, Qin Zhu, Fengnan Li et al.ACL 2026
- TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question AnsweringYushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang et al.ICCV 2023 · 400 citations
- Who Evaluates the Evaluations? Objectively Scoring Text-to-Image Prompt Coherence Metrics with T2IScoreScore (TS2)Michael Saxon, Fatima Jahara, Mahsa Khoshnoodi, Yujie Lu et al.NeurIPS 2024 · 19 citations
