Probabilistic Prompt Adaptation for Unified Image Aesthetics and Quality Assessment
Takayuki Hara, Yuya Otsuka
Abstract
Recent advances in vision-language foundation models have enabled text-driven evaluation of image aesthetics and visual quality. However, existing models are typically optimized for fixed prompts or specific datasets, limiting their adaptability to diverse evaluation criteria. This paper presents Probabilistic Prompt Adaptation (PPA), a unified probabilistic framework that flexibly predicts aesthetic and quality scores conditioned on arbitrary text prompts. PPA formulates score prediction as a mixture over prompts, dynamically estimating prompt suitability based on both image content and task context. By marginalizing over prompts pre-sampled from a large language model (LLM), it enables annotation-free training using only triplets of task, image, and score. Experiments across multiple IAA and IQA benchmarks demonstrate that PPA achieves consistent and perceptually aligned prompt-based scoring, allowing fine-grained control over evaluation semantics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 51fd2f5f-bacf-47ab-b6c6-f8ae0657bbcfBuilds on31
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- MUSIQ: Multi-scale Image Quality TransformerJunjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar et al.ICCV 2021 · 1,325 citations
- Exploring CLIP for Assessing the Look and Feel of ImagesJianyi Wang, Kelvin C. K. Chan, Chen Change LoyAAAI 2023 · 1,208 citations
Related papers
- Adaptive Prompt Learning for Blind Image Quality Assessment with Multi-modal Mixed-datasets TrainingYan Zhong, Xinping Zhao, Li Zhang, Xinyuan Song et al.ACM MM 2025
- CoFiVLA: Synergistic Coarse-Fine Vision-Language Alignment for Image Aesthetic AssessmentYuzhen Niu, Siling Chen, Yuzhong Chen, Fusheng Li et al.ACM MM 2025 · 1 citation
- Regression over Classification: Assessing Image Aesthetics via Multimodal Large Language ModelsXingyuan Ma, Shuai He, Anlong Ming, Haobin Zhong et al.AAAI 2026
- VILA: Learning Image Aesthetics from User Comments with Vision-Language PretrainingJunjie Ke, Keren Ye, Jiahui Yu, Yonghui Wu et al.CVPR 2023
- Attribute-Driven Multimodal Hierarchical Prompts for Image Aesthetic Quality AssessmentHancheng Zhu, Ju Shi, Zhiwen Shao, Rui Yao et al.ACM MM 2024 · 8 citations
