A-Bench: Are LMMs Masters at Evaluating AI-generated Images?
Zicheng Zhang, Haoning Wu, Chunyi Li, Yingjie Zhou, Wei Sun, Xiongkuo Min, Zijian Chen, Xiaohong Liu, Weisi Lin, Guangtao Zhai
Abstract
From Contradiction Overcome From Generative Distortion Assessment What is the most severe generative distortion? A. Incorrect structure of the handgun B. Blur due to low completion C. Incorrect structure of the woman's face D. Incorrect structure of the woman's hand (correct) GPT-4o Response: A Gemini 1.5 Pro Response : D Does the cactus contain soft and fluffy leaves? A. No B. Yes (correct) GPT-4o Response: A Gemini 1.5 Pro Response: A. No From Composition Identification What is partially covered by the mountain climber's backpacks? A. Climbing harnesses B. Boots lined up behind C. Ropes and carabiners D. The view of the mountain in the background (correct) GPT-4o Response: B. Boots lined up behind Gemini 1.5 Pro Response: B. Boots lined up behind Figure 1: Error cases from the A-Bench.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 42d2ba30-4bd9-4f1d-8cec-169f5db02b88Cited by top-tier papers15
- Veritas: Generalizable Deepfake Detection via Pattern-Aware ReasoningHao Tan, Jun Lan, Zichang Tan, Senyuan Shi et al.ICLR 2026 · 26 citations
- Agentic Retoucher for Text-To-Image GenerationShaocheng Shen, Jianfeng Liang, Chunlei Cai, Cong Geng et al.CVPR 2026 · 9 citations
- FakeXplain: AI-Generated Image Detection via Human-Aligned Grounded ReasoningYikun Ji, Yan Hong, Qi Fan, Jun Lan et al.ICLR 2026 · 9 citations
- Towards Explainable Fake Image Detection with Multi-Modal Large Language ModelsYikun Ji, Yan Hong, Jiahui Zhan, Haoxing Chen et al.ACM MM 2025 · 3 citations
- OBI-Bench: Can LMMs Aid in Study of Ancient Script on Oracle Bones?Zijian Chen, Tingzhu Chen, Wenjun Zhang, Guangtao ZhaiICLR 2025 · 3 citations
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
Related papers
- Eval3D: Interpretable and Fine-grained Evaluation for 3D GenerationShivam Duggal, Yushi Hu, Oscar Michel, Aniruddha Kembhavi et al.CVPR 2025
- Hash3D: Training-free Acceleration for 3D GenerationXingyi Yang, Songhua Liu, Xinchao WangCVPR 2025
- SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and TrainingJierun Chen, Dongting Hu, Xijie Huang, Huseyin Coskun et al.CVPR 2025
- ArtAdapter: Text-to-Image Style Transfer using Multi-Level Style Encoder and Explicit AdaptationDar-Yen Chen, Hamish Tennent, Ching-Wen HsuCVPR 2024
- Check, Locate, Rectify: A Training-Free Layout Calibration System for Text- to- Image GenerationBiao Gong, Siteng Huang, Yutong Feng, Shiwei Zhang et al.CVPR 2024
