SCIEval: Evaluating and Benchmarking the Faithfulness of Scientific Image Generation and Interpretation with Large Multimodal Models
Guanghui Ye, Huan Zhao, Zhixue Zhao, Tengfei Ma, Kehan Wang, Steffen Eger, Zhihua Jiang
摘要
Scientific images often require accurate numerical representations and correct object attributes. However, current faithfulness metrics are primarily tailored toward photorealistic, real-life imagery, rendering them ill-suited for scientific image evaluation. To address this gap, we introduce a novel evaluation model, SCIEval (SCientific Image Evaluation), which aims to capture faithfulness through three key dimensions: (i) Relevance, measuring overall textimage correspondence; (ii) Accuracy, examining the technical details of scientific objects; and (iii) Explainability, which isolates unfaithful elements within the generated content. To address these dimensions, we curate a specialized dataset of scientific text-image pairs to train three evaluation modules. For the Relevance and Accuracy modules, we propose a CLIP-based strategy that enhances scientific image perception through intra-and cross-modal contrastive learning. Concurrently, the Explainability module is developed by fine-tuning a high-performance Large Multimodal Model (LMM) using supervised rationale signals. Finally, we present SCIEval-Bench, a human-annotated evaluation benchmark consisting of 3,000 samples for scientific textto-image and 3,000 samples for scientific image captioning. Extensive experiments on SCIEval-Bench demonstrate that our SCIEval model is significantly more reliable than 24 competing models-including GPT-4o-exhibiting a superior correlation with human judgments. Project page: https://SCIEval.github.io
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras 等EMNLP 2021 · 被引用 937 次
相关 Paper
- ScImage: How good are multimodal large language models at scientific text-to-image generation?Leixin Zhang, Steffen Eger, Yinjie Cheng, Weihe Zhai 等ICLR 2025
- SciEval: A Multi-Level Large Language Model Evaluation Benchmark for Scientific ResearchLiangtai Sun, Yang Han, Zihan Zhao, Da Ma 等AAAI 2024 · 被引用 150 次
- Automatic and Reliable Evaluation for Academic Caption-to-Figure Generation with LMMsGuanghui Ye, Huan Zhao, Qin Zhu, Fengnan Li 等ACL 2026
- TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question AnsweringYushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang 等ICCV 2023 · 被引用 400 次
- Who Evaluates the Evaluations? Objectively Scoring Text-to-Image Prompt Coherence Metrics with T2IScoreScore (TS2)Michael Saxon, Fatima Jahara, Mahsa Khoshnoodi, Yujie Lu 等NeurIPS 2024 · 被引用 19 次
