QPT-V2: Masked Image Modeling Advances Visual Scoring
Qizhi Xie, Kun Yuan, Yunpeng Qu, Mingda Wu, Ming Sun, Chao Zhou, Jihong Zhu
摘要
Quality assessment and aesthetics assessment aim to evaluate the perceived quality and aesthetics of visual content. Current learning-based methods suffer greatly from the scarcity of labeled data and usually perform sub-optimally in terms of generalization. Although masked image modeling (MIM) has achieved noteworthy advancements across various high-level tasks (e.g., classification, detection etc.). In this work, we take on a novel perspective to investigate its capabilities in terms of quality- and aesthetics-awareness. To this end, we propose Quality- and aesthetics-aware pretraining (QPT V2), the first pretraining framework based on MIM that offers a unified solution to quality and aesthetics assessment. To perceive the high-level semantics and fine-grained details, pretraining data is curated. To comprehensively encompass quality- and aesthetics-related factors, degradation is introduced. To capture multi-scale quality and aesthetic information, model structure is modified. Extensive experimental results on 11 downstream benchmarks clearly show the superior performance of QPT V2 in comparison with current state-of-the-art approaches and other pretraining paradigms. Code and models will be released at https://github.com/KeiChiTse/QPT-V2.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Learning Disentangled Representations for Perceptual Point Cloud Quality Assessment via Mutual Information MinimizationZiyu Shan, Yujie Zhang, Yipeng Liu, Yiling XuNeurIPS 2024 · 被引用 7 次
- Plug-and-Play Tri-Branch Invertible Block for Image RescalingJingwei Bao, Jinhua Hao, Pengcheng Xu, Ming Sun 等AAAI 2025
- KVQ: Boosting Video Quality Assessment via Saliency-guided Local PerceptionYunpeng Qu, Kun Yuan, Qizhi Xie, Ming Sun 等CVPR 2025
- Score2Instruct: Scaling Up Video Quality-Centric Instructions via Automated Dimension ScoringQizhi Xie, Kun Yuan, Yunpeng Qu, Jiachao Gong 等CVPR 2026
它引用的顶会 Paper32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 被引用 2,336 次
相关 Paper
- Global Patch-wise Attention is Masterful Facilitator for Masked Image ModelingGongli Xi, Ye Tian, Mengyu Yang, Lanshan Zhang 等ACM MM 2024 · 被引用 1 次
- Attribute-Driven Multimodal Hierarchical Prompts for Image Aesthetic Quality AssessmentHancheng Zhu, Ju Shi, Zhiwen Shao, Rui Yao 等ACM MM 2024 · 被引用 8 次
- The Devil Is in the Frequency: Geminated Gestalt Autoencoder for Self-Supervised Visual Pre-trainingHao Liu, Xinghua Jiang, Xin Li, Antai Guo 等AAAI 2023 · 被引用 45 次
- Harmonizing Visual Representations for Unified Multimodal Understanding and GenerationSize Wu, Wenwei Zhang, Lumin Xu, Sheng Jin 等ICCV 2025 · 被引用 3 次
- Quality-aware Pretrained Models for Blind Image Quality AssessmentKai Zhao, Kun Yuan, Ming Sun, Mading Li 等CVPR 2023
