QPT-V2: Masked Image Modeling Advances Visual Scoring
Qizhi Xie, Kun Yuan, Yunpeng Qu, Mingda Wu, Ming Sun, Chao Zhou, Jihong Zhu
Abstract
Quality assessment and aesthetics assessment aim to evaluate the perceived quality and aesthetics of visual content. Current learning-based methods suffer greatly from the scarcity of labeled data and usually perform sub-optimally in terms of generalization. Although masked image modeling (MIM) has achieved noteworthy advancements across various high-level tasks (e.g., classification, detection etc.). In this work, we take on a novel perspective to investigate its capabilities in terms of quality- and aesthetics-awareness. To this end, we propose Quality- and aesthetics-aware pretraining (QPT V2), the first pretraining framework based on MIM that offers a unified solution to quality and aesthetics assessment. To perceive the high-level semantics and fine-grained details, pretraining data is curated. To comprehensively encompass quality- and aesthetics-related factors, degradation is introduced. To capture multi-scale quality and aesthetic information, model structure is modified. Extensive experimental results on 11 downstream benchmarks clearly show the superior performance of QPT V2 in comparison with current state-of-the-art approaches and other pretraining paradigms. Code and models will be released at https://github.com/KeiChiTse/QPT-V2.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1a34b76d-bd14-4ea8-88ea-1ee922be8c87Cited by top-tier papers4
- Learning Disentangled Representations for Perceptual Point Cloud Quality Assessment via Mutual Information MinimizationZiyu Shan, Yujie Zhang, Yipeng Liu, Yiling XuNeurIPS 2024 · 7 citations
- Plug-and-Play Tri-Branch Invertible Block for Image RescalingJingwei Bao, Jinhua Hao, Pengcheng Xu, Ming Sun et al.AAAI 2025
- KVQ: Boosting Video Quality Assessment via Saliency-guided Local PerceptionYunpeng Qu, Kun Yuan, Qizhi Xie, Ming Sun et al.CVPR 2025
- Score2Instruct: Scaling Up Video Quality-Centric Instructions via Automated Dimension ScoringQizhi Xie, Kun Yuan, Yunpeng Qu, Jiachao Gong et al.CVPR 2026
Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
Related papers
- Global Patch-wise Attention is Masterful Facilitator for Masked Image ModelingGongli Xi, Ye Tian, Mengyu Yang, Lanshan Zhang et al.ACM MM 2024 · 1 citation
- Attribute-Driven Multimodal Hierarchical Prompts for Image Aesthetic Quality AssessmentHancheng Zhu, Ju Shi, Zhiwen Shao, Rui Yao et al.ACM MM 2024 · 8 citations
- The Devil Is in the Frequency: Geminated Gestalt Autoencoder for Self-Supervised Visual Pre-trainingHao Liu, Xinghua Jiang, Xin Li, Antai Guo et al.AAAI 2023 · 45 citations
- Harmonizing Visual Representations for Unified Multimodal Understanding and GenerationSize Wu, Wenwei Zhang, Lumin Xu, Sheng Jin et al.ICCV 2025 · 3 citations
- Quality-aware Pretrained Models for Blind Image Quality AssessmentKai Zhao, Kun Yuan, Ming Sun, Mading Li et al.CVPR 2023
