ACL2026
Automatic and Reliable Evaluation for Academic Caption-to-Figure Generation with LMMs
Guanghui Ye, Huan Zhao, Qin Zhu, Fengnan Li, Jiaqi Li, Yixian Shen, Zhonghao Ren, Zhihua Jiang
摘要
Existing datasets for evaluating text-to-image generation focus mostly on real-life images, which poses challenges for assessing academic figure generation given real scientific captions, which is a hot topic in AI for Science. To fill the gap, we propose HE4AFG, a novel dataset which first provides a Holistic Evaluation for Academic caption-to-Figure Generation (AFG). Specifically, HE4AFG collects real figure captions from 8 scientific domains and finally generates 3,900 evaluation samples (particularly, including multi-panel figures) using 5 mainstream large multimodal models (LMMs). For each sample, we provide highquality human ratings in terms of three aspects-scientific aesthetic (SA), topic relevance (TR), and attribute correctness (AC). Moreover, we present two trainable models: (1) HE4AFG-E, an automated Evaluation model for AFG, which generates aspect-aware training examples and then use them to train three aspect-specific evaluation modules via contrastive learning; (2) HE4AFG-R, an automated Refinement model, which generates and utilizes feedback on the quality of the figures (e.g., unfaithful elements) to continuously improve AFG. Extensive experiments on HE4AFG demonstrate the effectiveness and performance advantages of our models. * * Corresponding authors. C: Seven green dashed lines do not intersect one another and maintain equal spacing. They pass through a gray plane and form identical angles with it. (b) Academic Caption-to-Figure (a) General Text-to-Image score: 5.0 Relevance: 3.5 Aesthetic: 4.0 Real-life Image (I) Previous datasets T: The bright and tidy exterior of the FamilyMart convenience store, with lights shining brightly inside. Academic Figure (F) (C, F) HE4AFG (ours) (T, I) Correctness: 3.0 Dataset Size Human-Annotated? Tasks SciCap (Hsu et al., 2021) 2M image-caption pairs No image captioning Multi-modal ArXiv (Li et al., 2024) 6.4M images and 3.9M captions No image captioning; question-answering SciFIBench (Roberts et al., 2024) 2,000 image-caption pairs No image captioning; question-answering CharXiv (Wang et al., 2024) 2,323 images No question-answering DaTikZ (Belouadi et al., 2024) 120,000 text-image pairs No text-to-image VGBench (Zou et al., 2024) 4,279 / 5,845 text-image pairs No question-answering / text-to-image ScImage (Zhang et al., 2025a) 3,000 images and 404 prompts Yes text-to-image (evaluation) Science-T2I (Li et al., 2025) 20,000 images and 9,000 prompts Yes text-to-image (hallucination detection) HE4AFG (ours) 3,900 images and 650 captions Yes caption-to-figure (evaluation)