HEIE: MLLM-Based Hierarchical Explainable AIGC Image Implausibility Evaluator
Fan Yang, Ru Zhen, Jianing Wang, Yanhao Zhang, Haoxiang Chen, Haonan Lu, Sicheng Zhao, Guiguang Ding
Abstract
AIGC images are prevalent across various fields, yet they frequently suffer from quality issues like artifacts and unnatural textures. Specialized models aim to predict defect region heatmaps but face two primary challenges: (1) lack of explainability, failing to provide reasons and analyses for subtle defects, and (2) inability to leverage common sense and logical reasoning, leading to poor generalization. Multimodal large language models (MLLMs) promise better comprehension and reasoning but face their own challenges: (1) difficulty in fine-grained defect localization due to the limitations in capturing tiny details, and (2) constraints in providing pixel-wise outputs necessary for precise heatmap generation. To address these challenges, we propose HEIE: a novel MLLM-Based Hierarchical Explainable Image Implausibility Evaluator. We introduce the CoT-Driven Explainable Trinity Evaluator, which integrates heatmaps, scores, and explanation outputs, using CoT to decompose complex tasks into subtasks of increasing difficulty and enhance interpretability. Our Adaptive Hierarchical Implausibility Mapper synergizes low-level image features with high-level mapper tokens from LLMs, enabling precise local-to-global hierarchical heatmap predictions through an uncertainty-based adaptive token approach. Moreover, we propose a new dataset: Expl-AIGI-Eval, designed to facilitate interpretable implausibility evaluation of AIGC images. Our method demonstrates state-of-the-art performance through extensive experiments. Our project is at https://yfthu.github.io/HEIE/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Thinking with Frames: Generative Video Distortion Evaluation via Frame Reward ModelYuan Wang, Borui Liao, Huijuan Huang, Jinda Lu et al.CVPR 2026 · 5 citations
- Unveiling Perceptual Artifacts: A Fine-Grained Benchmark for Interpretable AI-Generated Image DetectionYao Xiao, Weiyan Chen, Jiahao Chen, Zijie Cao et al.ICLR 2026 · 4 citations
- Diffusion Epistemic Uncertainty with Asymmetric Learning for Diffusion-Generated Image DetectionYingsong Huang, Hui Guo, Jing Huang, Bing Bai et al.ICCV 2025 · 4 citations
- Aigi-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language ModelsZiyin Zhou, Yunpeng Luo, Yuanchen Wu, Ke Sun et al.ICCV 2025 · 3 citations
- Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated ImagesYikun Ji, Yan Hong, Bowen Deng, Jun Lan et al.CVPR 2026 · 3 citations
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Towards Explainable Fake Image Detection with Multi-Modal Large Language ModelsYikun Ji, Yan Hong, Jiahui Zhan, Haoxing Chen et al.ACM MM 2025 · 3 citations
- FakeXplain: AI-Generated Image Detection via Human-Aligned Grounded ReasoningYikun Ji, Yan Hong, Qi Fan, Jun Lan et al.ICLR 2026 · 9 citations
- Reasoning-Driven Anomaly Detection and Localization with Image-Level SupervisionYizhou Jin, Yuezhu Feng, Jinjin Zhang, Peng Wang et al.CVPR 2026 · 4 citations
- Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token GenerationRuoyu Chen, Xiaoqing Guo, Kangwei Liu, Siyuan Liang et al.CVPR 2026 · 19 citations
- FREAK: A Fine-grained Hallucination Evaluation Benchmark for Advanced MLLMsZhihan Yin, Jianxin Liang, Yueqian Wang, Yifeng Yao et al.ICLR 2026 · 6 citations
