Attribute-Driven Multimodal Hierarchical Prompts for Image Aesthetic Quality Assessment
Hancheng Zhu, Ju Shi, Zhiwen Shao, Rui Yao, Yong Zhou, Jiaqi Zhao, Leida Li
Abstract
Image Aesthetic Quality Assessment (IAQA) aims to simulate users' visual perception to judge the aesthetic quality of images. In social media, users' aesthetic experiences are often reflected in their textual comments regarding the aesthetic attributes of images. To fully explore the attribute information perceived by users for evaluating image aesthetic quality, this paper proposes an image aesthetic quality assessment method based on attribute-driven multimodal hierarchical prompts. Unlike existing IAQA methods that utilize multimodal pre-training or straightforward prompts for model learning, the proposed method leverages attribute comments and quality-level text templates to hierarchically learn the aesthetic attributes and quality of images. Specifically, we first leverage users' aesthetic attribute comments to perform prompt learning on images. The learned attribute-driven multimodal features can comprehensively capture the semantic information of image aesthetic attributes perceived by users. Then, we construct text templates for different aesthetic quality levels to further facilitate prompt learning through semantic information related to the aesthetic quality of images. The proposed method can explicitly simulate users' aesthetic judgment of images to obtain more precise aesthetic quality. Experimental results demonstrate that the proposed IAQA method based on hierarchical prompts outperforms existing methods significantly on multiple IAQA databases. Our source code is public at https://github.com/GitHub-Ju/AMHP.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 2917ffbb-c2de-4608-bf8b-b1351f9b1778Cited by top-tier papers1
Ask how each one uses itRelated papers
- VILA: Learning Image Aesthetics from User Comments with Vision-Language PretrainingJunjie Ke, Keren Ye, Jiahui Yu, Yonghui Wu et al.CVPR 2023
- AesCLIP: Multi-Attribute Contrastive Learning for Image Aesthetics AssessmentXiangfei Sheng, Leida Li, Pengfei Chen, Jinjian Wu et al.ACM MM 2023 · 36 citations
- Personalized Image Aesthetics Assessment with Attribute-guided Fine-grained Feature RepresentationHancheng Zhu, Zhiwen Shao, Yong Zhou, Guangcheng Wang et al.ACM MM 2023 · 16 citations
- CoFiVLA: Synergistic Coarse-Fine Vision-Language Alignment for Image Aesthetic AssessmentYuzhen Niu, Siling Chen, Yuzhong Chen, Fusheng Li et al.ACM MM 2025 · 1 citation
- Aesthetically Relevant Image CaptioningZhipeng Zhong, Fei Zhou, Guoping QiuAAAI 2023 · 16 citations
