CoFiVLA: Synergistic Coarse-Fine Vision-Language Alignment for Image Aesthetic Assessment
Yuzhen Niu, Siling Chen, Yuzhong Chen, Fusheng Li, Rui Xu, Hui Da
摘要
Image aesthetic assessment (IAA) is challenging due to its subjective and diverse nature, making visual information alone insufficient. Existing methods using paired image and user comments, though effective, have limited practicality. Moreover, user comments contain both high- and weak-aesthetic-related textual information, thus directly integrating all textual information with visual information cannot guarantee the effectiveness. To address these issues, we propose a novel synergistic coarse-fine vision-language alignment (CoFiVLA) framework for IAA. It includes a CoFiVLA pretraining network and a CoFiVLA prediction network, aiming to effectively and comprehensively align and synergize image and text modalities at both coarse and fine granularities during training, while eliminating the dependence on user comments during inference. In the CoFiVLA pretraining network, the coarse- and fine-grained vision-language alignment branches work together, aligning the visual features with the high-aesthetic-related textual information and all textual information, respectively. In the coarse-grained alignment branch, we innovatively propose employing the large language model LLaMA to construct an aesthetic summary dataset, which extracts high-aesthetic-related text from user comments. Furthermore, in the CoFiVLA prediction network, we first extract features corresponding to learnable prompts of different aesthetic quality categories based on the CoFiVLA pretrained model, then fuse visual features with these learnable textual features. Thus we achieve aesthetic quality semantics embedded image aesthetic representations for effective IAA without requiring user comments during inference. Our aesthetic summary dataset and source code are available at https://github.com/lifusheng-chn/CoFiVLA.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- VILA: Learning Image Aesthetics from User Comments with Vision-Language PretrainingJunjie Ke, Keren Ye, Jiahui Yu, Yonghui Wu 等CVPR 2023
- Advancing Comprehensive Aesthetic Insight with Multi-Scale Text-Guided Self-Supervised LearningYuti Liu, Shice Liu, Junyuan Gao, Peng-Tao Jiang 等AAAI 2025
- Attribute-Driven Multimodal Hierarchical Prompts for Image Aesthetic Quality AssessmentHancheng Zhu, Ju Shi, Zhiwen Shao, Rui Yao 等ACM MM 2024 · 被引用 8 次
- Probabilistic Prompt Adaptation for Unified Image Aesthetics and Quality AssessmentTakayuki Hara, Yuya OtsukaCVPR 2026
- AesCLIP: Multi-Attribute Contrastive Learning for Image Aesthetics AssessmentXiangfei Sheng, Leida Li, Pengfei Chen, Jinjian Wu 等ACM MM 2023 · 被引用 36 次
