Lune

ACL2026顶会

Automatic and Reliable Evaluation for Academic Caption-to-Figure Generation with LMMs

Guanghui Ye, Huan Zhao, Qin Zhu, Fengnan Li, Jiaqi Li, Yixian Shen, Zhonghao Ren, Zhihua Jiang

2026年份

摘要

Existing datasets for evaluating text-to-image generation focus mostly on real-life images, which poses challenges for assessing academic figure generation given real scientific captions, which is a hot topic in AI for Science. To fill the gap, we propose HE4AFG, a novel dataset which first provides a Holistic Evaluation for Academic caption-to-Figure Generation (AFG). Specifically, HE4AFG collects real figure captions from 8 scientific domains and finally generates 3,900 evaluation samples (particularly, including multi-panel figures) using 5 mainstream large multimodal models (LMMs). For each sample, we provide highquality human ratings in terms of three aspects-scientific aesthetic (SA), topic relevance (TR), and attribute correctness (AC). Moreover, we present two trainable models: (1) HE4AFG-E, an automated Evaluation model for AFG, which generates aspect-aware training examples and then use them to train three aspect-specific evaluation modules via contrastive learning; (2) HE4AFG-R, an automated Refinement model, which generates and utilizes feedback on the quality of the figures (e.g., unfaithful elements) to continuously improve AFG. Extensive experiments on HE4AFG demonstrate the effectiveness and performance advantages of our models. * * Corresponding authors. C: Seven green dashed lines do not intersect one another and maintain equal spacing. They pass through a gray plane and form identical angles with it. (b) Academic Caption-to-Figure (a) General Text-to-Image score: 5.0 Relevance: 3.5 Aesthetic: 4.0 Real-life Image (I) Previous datasets T: The bright and tidy exterior of the FamilyMart convenience store, with lights shining brightly inside. Academic Figure (F) (C, F) HE4AFG (ours) (T, I) Correctness: 3.0 Dataset Size Human-Annotated? Tasks SciCap (Hsu et al., 2021) 2M image-caption pairs No image captioning Multi-modal ArXiv (Li et al., 2024) 6.4M images and 3.9M captions No image captioning; question-answering SciFIBench (Roberts et al., 2024) 2,000 image-caption pairs No image captioning; question-answering CharXiv (Wang et al., 2024) 2,323 images No question-answering DaTikZ (Belouadi et al., 2024) 120,000 text-image pairs No text-to-image VGBench (Zou et al., 2024) 4,279 / 5,845 text-image pairs No question-answering / text-to-image ScImage (Zhang et al., 2025a) 3,000 images and 404 prompts Yes text-to-image (evaluation) Science-T2I (Li et al., 2025) 20,000 images and 9,000 prompts Yes text-to-image (hallucination detection) HE4AFG (ours) 3,900 images and 650 captions Yes caption-to-figure (evaluation)

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper23

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖