Automatic and Reliable Evaluation for Academic Caption-to-Figure Generation with LMMs
Guanghui Ye, Huan Zhao, Qin Zhu, Fengnan Li, Jiaqi Li, Yixian Shen, Zhonghao Ren, Zhihua Jiang
Abstract
Existing datasets for evaluating text-to-image generation focus mostly on real-life images, which poses challenges for assessing academic figure generation given real scientific captions, which is a hot topic in AI for Science. To fill the gap, we propose HE4AFG, a novel dataset which first provides a Holistic Evaluation for Academic caption-to-Figure Generation (AFG). Specifically, HE4AFG collects real figure captions from 8 scientific domains and finally generates 3,900 evaluation samples (particularly, including multi-panel figures) using 5 mainstream large multimodal models (LMMs). For each sample, we provide highquality human ratings in terms of three aspects-scientific aesthetic (SA), topic relevance (TR), and attribute correctness (AC). Moreover, we present two trainable models: (1) HE4AFG-E, an automated Evaluation model for AFG, which generates aspect-aware training examples and then use them to train three aspect-specific evaluation modules via contrastive learning; (2) HE4AFG-R, an automated Refinement model, which generates and utilizes feedback on the quality of the figures (e.g., unfaithful elements) to continuously improve AFG. Extensive experiments on HE4AFG demonstrate the effectiveness and performance advantages of our models. * * Corresponding authors. C: Seven green dashed lines do not intersect one another and maintain equal spacing. They pass through a gray plane and form identical angles with it. (b) Academic Caption-to-Figure (a) General Text-to-Image score: 5.0 Relevance: 3.5 Aesthetic: 4.0 Real-life Image (I) Previous datasets T: The bright and tidy exterior of the FamilyMart convenience store, with lights shining brightly inside. Academic Figure (F) (C, F) HE4AFG (ours) (T, I) Correctness: 3.0 Dataset Size Human-Annotated? Tasks SciCap (Hsu et al., 2021) 2M image-caption pairs No image captioning Multi-modal ArXiv (Li et al., 2024) 6.4M images and 3.9M captions No image captioning; question-answering SciFIBench (Roberts et al., 2024) 2,000 image-caption pairs No image captioning; question-answering CharXiv (Wang et al., 2024) 2,323 images No question-answering DaTikZ (Belouadi et al., 2024) 120,000 text-image pairs No text-to-image VGBench (Zou et al., 2024) 4,279 / 5,845 text-image pairs No question-answering / text-to-image ScImage (Zhang et al., 2025a) 3,000 images and 404 prompts Yes text-to-image (evaluation) Science-T2I (Li et al., 2025) 20,000 images and 9,000 prompts Yes text-to-image (hallucination detection) HE4AFG (ours) 3,900 images and 650 captions Yes caption-to-figure (evaluation)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext edbb6d77-9c3a-4f80-97dc-21eb01856bf5Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image GenerationYuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana et al.NeurIPS 2023 · 1,192 citations
Related papers
- SCIEval: Evaluating and Benchmarking the Faithfulness of Scientific Image Generation and Interpretation with Large Multimodal ModelsGuanghui Ye, Huan Zhao, Zhixue Zhao, Tengfei Ma et al.CVPR 2026
- What Makes a Good Generated Image? Investigating Human and Multimodal LLM Image Preference AlignmentRishab Parthasarathy, Jasmine Collins, Cory StephensonAAAI 2026
- Learning Multi-Dimensional Human Preference for Text-to-Image GenerationSixian Zhang, Bohan Wang, Junqiang Wu, Yan Li et al.CVPR 2024 · 15 citations
- D-Judge: How Far Are We? Assessing the Discrepancies Between AI-synthesized and Natural Images through Multimodal GuidanceRenyang Liu, Ziyu Lyu, Wei Zhou, See-Kiong NgACM MM 2025
- Too Vivid to Be Real? Benchmarking and Calibrating Generative Color FidelityZhengyao Fang, Zexi Jia, Yijia Zhong, Pengcheng Luo et al.CVPR 2026 · 1 citation
