CoFiVLA: Synergistic Coarse-Fine Vision-Language Alignment for Image Aesthetic Assessment
Yuzhen Niu, Siling Chen, Yuzhong Chen, Fusheng Li, Rui Xu, Hui Da
Abstract
Image aesthetic assessment (IAA) is challenging due to its subjective and diverse nature, making visual information alone insufficient. Existing methods using paired image and user comments, though effective, have limited practicality. Moreover, user comments contain both high- and weak-aesthetic-related textual information, thus directly integrating all textual information with visual information cannot guarantee the effectiveness. To address these issues, we propose a novel synergistic coarse-fine vision-language alignment (CoFiVLA) framework for IAA. It includes a CoFiVLA pretraining network and a CoFiVLA prediction network, aiming to effectively and comprehensively align and synergize image and text modalities at both coarse and fine granularities during training, while eliminating the dependence on user comments during inference. In the CoFiVLA pretraining network, the coarse- and fine-grained vision-language alignment branches work together, aligning the visual features with the high-aesthetic-related textual information and all textual information, respectively. In the coarse-grained alignment branch, we innovatively propose employing the large language model LLaMA to construct an aesthetic summary dataset, which extracts high-aesthetic-related text from user comments. Furthermore, in the CoFiVLA prediction network, we first extract features corresponding to learnable prompts of different aesthetic quality categories based on the CoFiVLA pretrained model, then fuse visual features with these learnable textual features. Thus we achieve aesthetic quality semantics embedded image aesthetic representations for effective IAA without requiring user comments during inference. Our aesthetic summary dataset and source code are available at https://github.com/lifusheng-chn/CoFiVLA.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get fab12a16-5ab9-4ca3-8c7f-46b2dc5ce2f2Related papers
- VILA: Learning Image Aesthetics from User Comments with Vision-Language PretrainingJunjie Ke, Keren Ye, Jiahui Yu, Yonghui Wu et al.CVPR 2023
- Advancing Comprehensive Aesthetic Insight with Multi-Scale Text-Guided Self-Supervised LearningYuti Liu, Shice Liu, Junyuan Gao, Peng-Tao Jiang et al.AAAI 2025
- Attribute-Driven Multimodal Hierarchical Prompts for Image Aesthetic Quality AssessmentHancheng Zhu, Ju Shi, Zhiwen Shao, Rui Yao et al.ACM MM 2024 · 8 citations
- Probabilistic Prompt Adaptation for Unified Image Aesthetics and Quality AssessmentTakayuki Hara, Yuya OtsukaCVPR 2026
- AesCLIP: Multi-Attribute Contrastive Learning for Image Aesthetics AssessmentXiangfei Sheng, Leida Li, Pengfei Chen, Jinjian Wu et al.ACM MM 2023 · 36 citations
