Bridging Cognitive Gap: Hierarchical Description Learning for Artistic Image Aesthetics Assessment
Henglin Liu, Nisha Huang, Chang Liu, Jiangpeng Yan, Huijuan Huang, Jixuan Ying, Tong-Yee Lee, Pengfei Wan, Xiangyang Ji
Abstract
The aesthetic quality assessment task is crucial for developing a human-aligned quantitative evaluation system for AIGC. However, its inherently complex nature-spanning visual perception, cognition, and emotion-poses fundamental challenges. Although aesthetic descriptions offer a viable representation of this complexity, two critical challenges persist: (1) data scarcity and imbalance: existing dataset overly focuses on visual perception and neglects deeper dimensions due to the expensive manual annotation; and (2) model fragmentation: current visual networks isolate aesthetic attributes with multi-branch encoder, while multimodal methods represented by contrastive learning struggle to effectively process long-form textual descriptions. To resolve challenge (1), we first present the Refined Aesthetic Description (RAD) dataset, a large-scale (70k), multi-dimensional structured dataset, generated via an iterative pipeline without heavy annotation costs and easy to scale. To address challenge (2), we propose ArtQuant, an aesthetics assessment framework for artistic image which not only couple isolated aesthetic dimensions through joint description generation, but also better model long-text semantics with the help of LLM decoders. Besides, theoretical analysis confirms this symbiosis: RAD's semantic adequacy (data) and generation paradigm (model) collectively minimize prediction entropy, providing mathematical grounding for the framework. Our approach achieves state-of-the-art performance on several datasets while requiring only 33% of conventional training epochs, narrowing the cognitive gap between artistic image and aesthetic judgment. The code is available at: https://github.com/Henglin-Liu/ArtQuant .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 86a52672-ea33-4383-a4c1-bb3d7f1b53b3Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined LevelsHaoning Wu, Zicheng Zhang, Weixia Zhang, Chaofeng Chen et al.ICML 2024 · 499 citations
- Large Language Models for Data Annotation and Synthesis: A SurveyZhen Tan, Dawei Li, Song Wang, Alimohammad Beigi et al.EMNLP 2024 · 119 citations
- GalleryGPT: Analyzing Paintings with Large Multimodal ModelsYi Bin, Wenhao Shi, Yujuan Ding, Zhiqiang Hu et al.ACM MM 2024 · 35 citations
- MaTe: Images are All You Need for Material Transfer via Diffusion TransformerNisha Huang, Henglin Liu, Yizhou Lin, Kaer Huang et al.ICCV 2025 · 2 citations
- Towards Artistic Image Aesthetics Assessment: a Large-scale Dataset and a New MethodRan Yi, Haoyuan Tian, Zhihao Gu, Yu-Kun Lai et al.CVPR 2023
Related papers
- ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level UnderstandingShuo Cao, Nan Ma, Jiayang Li, Xiaohui Li et al.CVPR 2026 · 38 citations
- D-Judge: How Far Are We? Assessing the Discrepancies Between AI-synthesized and Natural Images through Multimodal GuidanceRenyang Liu, Ziyu Lyu, Wei Zhou, See-Kiong NgACM MM 2025
- AesCLIP: Multi-Attribute Contrastive Learning for Image Aesthetics AssessmentXiangfei Sheng, Leida Li, Pengfei Chen, Jinjian Wu et al.ACM MM 2023 · 36 citations
- VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality EvaluationLongteng Jiang, Dandan Zheng, Qianqian Qiao, Heng Huang et al.CVPR 2026 · 2 citations
- A³: Towards Advertising Aesthetic AssessmentKaiyuan Ji, Yixuan Gao, Lu Sun, Yushuo Zheng et al.CVPR 2026
