Automatic and Reliable Evaluation for Academic Caption-to-Figure Generation with LMMs
Guanghui Ye, Huan Zhao, Qin Zhu, Fengnan Li, Jiaqi Li, Yixian Shen, Zhonghao Ren, Zhihua Jiang
摘要
Existing datasets for evaluating text-to-image generation focus mostly on real-life images, which poses challenges for assessing academic figure generation given real scientific captions, which is a hot topic in AI for Science. To fill the gap, we propose HE4AFG, a novel dataset which first provides a Holistic Evaluation for Academic caption-to-Figure Generation (AFG). Specifically, HE4AFG collects real figure captions from 8 scientific domains and finally generates 3,900 evaluation samples (particularly, including multi-panel figures) using 5 mainstream large multimodal models (LMMs). For each sample, we provide highquality human ratings in terms of three aspects-scientific aesthetic (SA), topic relevance (TR), and attribute correctness (AC). Moreover, we present two trainable models: (1) HE4AFG-E, an automated Evaluation model for AFG, which generates aspect-aware training examples and then use them to train three aspect-specific evaluation modules via contrastive learning; (2) HE4AFG-R, an automated Refinement model, which generates and utilizes feedback on the quality of the figures (e.g., unfaithful elements) to continuously improve AFG. Extensive experiments on HE4AFG demonstrate the effectiveness and performance advantages of our models. * * Corresponding authors. C: Seven green dashed lines do not intersect one another and maintain equal spacing. They pass through a gray plane and form identical angles with it. (b) Academic Caption-to-Figure (a) General Text-to-Image score: 5.0 Relevance: 3.5 Aesthetic: 4.0 Real-life Image (I) Previous datasets T: The bright and tidy exterior of the FamilyMart convenience store, with lights shining brightly inside. Academic Figure (F) (C, F) HE4AFG (ours) (T, I) Correctness: 3.0 Dataset Size Human-Annotated? Tasks SciCap (Hsu et al., 2021) 2M image-caption pairs No image captioning Multi-modal ArXiv (Li et al., 2024) 6.4M images and 3.9M captions No image captioning; question-answering SciFIBench (Roberts et al., 2024) 2,000 image-caption pairs No image captioning; question-answering CharXiv (Wang et al., 2024) 2,323 images No question-answering DaTikZ (Belouadi et al., 2024) 120,000 text-image pairs No text-to-image VGBench (Zou et al., 2024) 4,279 / 5,845 text-image pairs No question-answering / text-to-image ScImage (Zhang et al., 2025a) 3,000 images and 404 prompts Yes text-to-image (evaluation) Science-T2I (Li et al., 2025) 20,000 images and 9,000 prompts Yes text-to-image (hallucination detection) HE4AFG (ours) 3,900 images and 650 captions Yes caption-to-figure (evaluation)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image GenerationYuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana 等NeurIPS 2023 · 被引用 1,192 次
相关 Paper
- SCIEval: Evaluating and Benchmarking the Faithfulness of Scientific Image Generation and Interpretation with Large Multimodal ModelsGuanghui Ye, Huan Zhao, Zhixue Zhao, Tengfei Ma 等CVPR 2026
- What Makes a Good Generated Image? Investigating Human and Multimodal LLM Image Preference AlignmentRishab Parthasarathy, Jasmine Collins, Cory StephensonAAAI 2026
- Learning Multi-Dimensional Human Preference for Text-to-Image GenerationSixian Zhang, Bohan Wang, Junqiang Wu, Yan Li 等CVPR 2024 · 被引用 15 次
- D-Judge: How Far Are We? Assessing the Discrepancies Between AI-synthesized and Natural Images through Multimodal GuidanceRenyang Liu, Ziyu Lyu, Wei Zhou, See-Kiong NgACM MM 2025
- Too Vivid to Be Real? Benchmarking and Calibrating Generative Color FidelityZhengyao Fang, Zexi Jia, Yijia Zhong, Pengcheng Luo 等CVPR 2026 · 被引用 1 次
