Regression over Classification: Assessing Image Aesthetics via Multimodal Large Language Models
Xingyuan Ma, Shuai He, Anlong Ming, Haobin Zhong, Huadong Ma
Abstract
Image Aesthetics Assessment (IAA) evaluates visual quality through user-centered perceptual analysis and can guide various applications. Recent advances in Multimodal Large Language Models (MLLMs) have sparked interest in adapting them for IAA. However, two critical limitations persist in applying MLLMs to IAA: 1) the tokenization strategy leads to insensitivity to scores, and 2) the classification-based decoding mechanisms introduce score quantization errors. Current MLLM-based IAA methods treat the task as coarse rating classification followed by probability-to-score mapping, which loses fine-grained information. To address these challenges, we propose ROC4MLLM, offering complementary solutions from two perspectives:1) Representation: We separate scores from the word token space to avoid tokenizing scores as text. An independent position token bridges these spaces, improving the sensitivity of the model to score positions in text. 2) Computation: We apply distinct loss functions for text and score predictions to enhance the sensitivity of the model to score gradients. Decoupling scores from text ensures effective supervision while preventing interference between scores and text in the loss computation. Extensive experiments across five datasets demonstrate that ROC4MLLM achieves state-of-the-art performance without requiring additional training data. Additionally, its plug-and-play design ensures seamless integration with existing MLLMs, boosting their IAA performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- MUSIQ: Multi-scale Image Quality TransformerJunjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar et al.ICCV 2021 · 1,325 citations
- Thinking Image Color Aesthetics Assessment: Models, Datasets and BenchmarksShuai He, Anlong Ming, Yaqi Li, Jinyuan Sun et al.ICCV 2023 · 36 citations
- EAT: An Enhancer for Aesthetics-Oriented TransformersShuai He, Anlong Ming, Shuntian Zheng, Haobin Zhong et al.ACM MM 2023 · 29 citations
- AesMamba: Universal Image Aesthetic Assessment with State Space ModelsFei Gao, Yuhao Lin, Jiaqi Shi, Maoying Qiao et al.ACM MM 2024 · 12 citations
- VILA: Learning Image Aesthetics from User Comments with Vision-Language PretrainingJunjie Ke, Keren Ye, Jiahui Yu, Yonghui Wu et al.CVPR 2023
Related papers
- Revisiting MLLM Based Image Quality Assessment: Errors and RemedyZhenchen Tang, Songlin Yang, Bo Peng, Zichuan Wang et al.AAAI 2026 · 2 citations
- Advancing Comprehensive Aesthetic Insight with Multi-Scale Text-Guided Self-Supervised LearningYuti Liu, Shice Liu, Junyuan Gao, Peng-Tao Jiang et al.AAAI 2025
- Teaching Large Language Models to Regress Accurate Image Quality Scores Using Score DistributionZhiyuan You, Xin Cai, Jinjin Gu, Tianfan Xue et al.CVPR 2025
- Q-Insight: Understanding Image Quality via Visual Reinforcement LearningWeiqi Li, Xuanyu Zhang, Shijie Zhao, Yabin Zhang et al.NeurIPS 2025 · 117 citations
- Probabilistic Prompt Adaptation for Unified Image Aesthetics and Quality AssessmentTakayuki Hara, Yuya OtsukaCVPR 2026
