Capturing Co-existing Distortions in User-Generated Content for No-reference Video Quality Assessment
Kun Yuan, Zishang Kong, Chuanchuan Zheng, Ming Sun, Xing Wen
Abstract
Video Quality Assessment (VQA), which aims to predict the perceptual quality of a video, has attracted raising attention with the rapid development of streaming media technology, such as Facebook, TikTok, Kwai, and so on. Compared with other sequence-based visual tasks (e.g., action recognition), VQA faces two under-estimated challenges unresolved in User Generated Content (UGC) videos. First, it is not rare that several frames containing serious distortions (e.g., blocking, blurriness), can determine the perceptual quality of the whole video, while other sequence-based tasks require more frames of equal importance for representations. Second, the perceptual quality of a video exhibits a multi-distortion distribution, due to the differences in the duration and probability of occurrence for various distortions. In order to solve the above challenges, we propose Visual Quality Transformer (VQT) to extract quality-related sparse features more efficiently. Methodologically, a Sparse Temporal Attention (STA) is proposed to sample keyframes by analyzing the temporal correlation between frames, which reduces the computational complexity from 𝑂 (𝑇 2 ) to 𝑂 (𝑇 log𝑇 ). Structurally, a Multi-Pathway Temporal Network (MPTN) utilizes multiple STA modules with different degrees of sparsity in parallel, capturing coexisting distortions in a video. Experimentally, VQT demonstrates superior performance than many state-of-the-art methods in three public no-reference VQA datasets. Furthermore, VQT shows better performance in four full-reference VQA datasets against widelyadopted industrial algorithms (i.e., VMAF and AVQT).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 958f7684-d962-4b80-af53-b33bc269c7d5Cited by top-tier papers9
- KVQ: Kwai Video Quality Assessment for Short-form VideosYiting Lu, Xin Li, Yajing Pei, Kun Yuan et al.CVPR 2024 · 32 citations
- CPGA: Coding Priors-Guided Aggregation Network for Compressed Video Quality EnhancementQiang Zhu, Jinhua Hao, Yukang Ding, Yu Liu et al.CVPR 2024 · 13 citations
- FVQ: A Large-Scale Dataset and an LMM-based Method for Face Video Quality AssessmentSijing Wu, Yunhao Li, Ziwen Xu, Yixuan Gao et al.ACM MM 2025 · 8 citations
- PTM-VQA: Efficient Video Quality Assessment Leveraging Diverse PreTrained Models from the WildKun Yuan, Hongbo Liu, Mading Li, Muyi Sun et al.CVPR 2024 · 8 citations
- QPT-V2: Masked Image Modeling Advances Visual ScoringQizhi Xie, Kun Yuan, Yunpeng Qu, Mingda Wu et al.ACM MM 2024 · 2 citations
Builds on15
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li et al.ICCV 2021 · 1,611 citations
Related papers
- A Deep Learning based No-reference Quality Assessment Model for UGC VideosWei Sun, Xiongkuo Min, Wei Lu, Guangtao ZhaiACM MM 2022 · 239 citations
- Patch-VQ: 'Patching Up' the Video Quality ProblemZhenqiang Ying, Maniratnam Mandal, Deepti Ghadiyaram, Alan C. BovikCVPR 2021
- Long Short-term Convolutional Transformer for No-Reference Video Quality AssessmentJunyong YouACM MM 2021 · 46 citations
- Perceptual Quality Assessment of Internet VideosJiahua Xu, Jing Li, Xingguang Zhou, Wei Zhou et al.ACM MM 2021 · 44 citations
- Blind Natural Video Quality Prediction via Statistical Temporal Features and Deep Spatial FeaturesJari Korhonen, Yicheng Su, Junyong YouACM MM 2020 · 88 citations
