Scaling and Masking: A New Paradigm of Data Sampling for Image and Video Quality Assessment
Yongxu Liu, Yinghui Quan, Guoyao Xiao, Aobo Li, Jinjian Wu
Abstract
Quality assessment of images and videos emphasizes both local details and global semantics, whereas general data sampling methods (e.g., resizing, cropping or grid-based fragment) fail to catch them simultaneously. To address the deficiency, current approaches have to adopt multi-branch models and take as input the multi-resolution data, which burdens the model complexity. In this work, instead of stacking up models, a more elegant data sampling method (named as SAMA, scaling and masking) is explored, which compacts both the local and global content in a regular input size. The basic idea is to scale the data into a pyramid first, and reduce the pyramid into a regular data dimension with a masking strategy. Benefiting from the spatial and temporal redundancy in images and videos, the processed data maintains the multi-scale characteristics with a regular input size, thus can be processed by a single-branch model. We verify the sampling method in image and video quality assessment. Experiments show that our sampling method can improve the performance of current single-branch models significantly, and achieves competitive performance to the multi-branch models without extra model complexity. The source code will be available at https://github.com/Sissuire/SAMA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c30395aa-160e-4407-b32f-773ffe07ce9cCited by top-tier papers3
- Highly Efficient No-reference 4K Video Quality Assessment with Full-Pixel Covering Sampling and Training StrategyXiaoheng Tan, Jiabin Zhang, Yuhui Quan, Jing Li et al.ACM MM 2024 · 2 citations
- MVQA: Mamba with Unified Sampling for Efficient Video Quality AssessmentYachun Mi, Yu Li, Weicheng Meng, Chaofeng Chen et al.ICCV 2025 · 1 citation
- Bridging the Perceptual Gap: Residual-Enhanced Downscaling and Manifold-Aware Perception Alignment Adaptation for NR-IQAYu Li, Zhengran Shen, Yachun Mi, Puchao Zhou et al.ICML 2026
Builds on18
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao et al.CVPR 2022 · 2,138 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- MUSIQ: Multi-scale Image Quality TransformerJunjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar et al.ICCV 2021 · 1,325 citations
- Exploring CLIP for Assessing the Look and Feel of ImagesJianyi Wang, Kelvin C. K. Chan, Chen Change LoyAAAI 2023 · 1,208 citations
Related papers
- A Deep Learning based No-reference Quality Assessment Model for UGC VideosWei Sun, Xiongkuo Min, Wei Lu, Guangtao ZhaiACM MM 2022 · 239 citations
- Modular Blind Video Quality AssessmentWen Wen, Mu Li, Yabin Zhang, Yiting Liao et al.CVPR 2024 · 23 citations
- Semantic-Aware and Quality-Aware Interaction Network for Blind Video Quality AssessmentJianjun Xiang, Yuanjie Dang, Peng Chen, Ronghua Liang et al.ACM MM 2024
- Domain-Specific Fusion Of Objective Video Quality MetricsAaron Chadha, Ioannis Katsavounidis, Ayan Kumar Bhunia, Cosmin Stejerean et al.ACM MM 2022 · 1 citation
- VSumMamba: Mamba Empowered Efficient Video Summarization with Multi-Scale Spatial-Temporal ModelingYamiao Ding, Tianrui Liu, Zhizhou Lu, Jun-Jie Huang et al.ACM MM 2025 · 1 citation
