AesMamba: Universal Image Aesthetic Assessment with State Space Models
Fei Gao, Yuhao Lin, Jiaqi Shi, Maoying Qiao, Nannan Wang
摘要
Image Aesthetic Assessment (IAA) aims to objectively predict the generic or personalized evaluations, of the aesthetic or fine-grained multi-attributes, based on visual or multimodal inputs. Previously, researchers have designed diverse and specialized methods, for specific IAA tasks, based on different input-output situations. Is it possible to design a universal IAA framework applicable for the whole IAA task taxonomy? In this paper, we explore this issue, and propose a modular IAA framework, dubbed AesMamba. Specially, we use the Visual State Space Model (VMamba), instead of CNNs or ViTs, to learn comprehensive representations of aesthetic-related attributes; because VMamba can efficiently achieve both global and local effective receptive fields. Afterward, a modal-adaptive module is used to automatically produce the integrated representations, conditioned on the type of input. In the prediction module, we propose a Multitask Balanced Adaptation (MBA) module, to boost task-specific features, with emphasis on the tail instances. Finally, we formulate the personalized IAA task as a multimodal learning problem, by converting a user's anonymous subject characters to a text prompt. This prompting strategy effectively employs the semantics of flexibly selected characters, for inferring individual preferences. AesMamba can be applied to diverse IAA tasks, through flexible combination of these modules. Extensive experiments on numerous datasets, demonstrate that AesMamba consistently achieves superior or competitive performance, on all IAA tasks, in comparison with previous SOTA methods. The code has been released at https://github.com/AiArt-Gao/AesMamba Github.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level UnderstandingShuo Cao, Nan Ma, Jiayang Li, Xiaohui Li 等CVPR 2026 · 被引用 38 次
- DiffoR: A Unified Continuous Generative Framework for Universal Ordinal RegressionHongxu Ma, Lin Wang, Chenghou Jin, Han Zhou 等KDD 2026 · 被引用 1 次
- MVQA: Mamba with Unified Sampling for Efficient Video Quality AssessmentYachun Mi, Yu Li, Weicheng Meng, Chaofeng Chen 等ICCV 2025 · 被引用 1 次
- Regression over Classification: Assessing Image Aesthetics via Multimodal Large Language ModelsXingyuan Ma, Shuai He, Anlong Ming, Haobin Zhong 等AAAI 2026
- GoR: A Unified and Extensible Generative Framework for Ordinal RegressionHongxu Ma, Han Zhou, Kai Tian, Xuefeng Zhang 等ICLR 2026
相关 Paper
- Personalized Image Aesthetics Assessment with Attribute-guided Fine-grained Feature RepresentationHancheng Zhu, Zhiwen Shao, Yong Zhou, Guangcheng Wang 等ACM MM 2023 · 被引用 16 次
- QMamba: On First Exploration of Vision Mamba for Image Quality AssessmentFengbin Guan, Xin Li, Zihao Yu, Yiting Lu 等ICML 2025
- Advancing Comprehensive Aesthetic Insight with Multi-Scale Text-Guided Self-Supervised LearningYuti Liu, Shice Liu, Junyuan Gao, Peng-Tao Jiang 等AAAI 2025
- AesCLIP: Multi-Attribute Contrastive Learning for Image Aesthetics AssessmentXiangfei Sheng, Leida Li, Pengfei Chen, Jinjian Wu 等ACM MM 2023 · 被引用 36 次
- Context-aware Attention Network for Predicting Image Aesthetic SubjectivityMunan Xu, Jia-Xing Zhong, Yurui Ren, Shan Liu 等ACM MM 2020 · 被引用 22 次
