Capturing Co-existing Distortions in User-Generated Content for No-reference Video Quality Assessment
Kun Yuan, Zishang Kong, Chuanchuan Zheng, Ming Sun, Xing Wen
摘要
Video Quality Assessment (VQA), which aims to predict the perceptual quality of a video, has attracted raising attention with the rapid development of streaming media technology, such as Facebook, TikTok, Kwai, and so on. Compared with other sequence-based visual tasks (e.g., action recognition), VQA faces two under-estimated challenges unresolved in User Generated Content (UGC) videos. First, it is not rare that several frames containing serious distortions (e.g., blocking, blurriness), can determine the perceptual quality of the whole video, while other sequence-based tasks require more frames of equal importance for representations. Second, the perceptual quality of a video exhibits a multi-distortion distribution, due to the differences in the duration and probability of occurrence for various distortions. In order to solve the above challenges, we propose Visual Quality Transformer (VQT) to extract quality-related sparse features more efficiently. Methodologically, a Sparse Temporal Attention (STA) is proposed to sample keyframes by analyzing the temporal correlation between frames, which reduces the computational complexity from 𝑂 (𝑇 2 ) to 𝑂 (𝑇 log𝑇 ). Structurally, a Multi-Pathway Temporal Network (MPTN) utilizes multiple STA modules with different degrees of sparsity in parallel, capturing coexisting distortions in a video. Experimentally, VQT demonstrates superior performance than many state-of-the-art methods in three public no-reference VQA datasets. Furthermore, VQT shows better performance in four full-reference VQA datasets against widelyadopted industrial algorithms (i.e., VMAF and AVQT).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- KVQ: Kwai Video Quality Assessment for Short-form VideosYiting Lu, Xin Li, Yajing Pei, Kun Yuan 等CVPR 2024 · 被引用 32 次
- CPGA: Coding Priors-Guided Aggregation Network for Compressed Video Quality EnhancementQiang Zhu, Jinhua Hao, Yukang Ding, Yu Liu 等CVPR 2024 · 被引用 13 次
- FVQ: A Large-Scale Dataset and an LMM-based Method for Face Video Quality AssessmentSijing Wu, Yunhao Li, Ziwen Xu, Yixuan Gao 等ACM MM 2025 · 被引用 8 次
- PTM-VQA: Efficient Video Quality Assessment Leveraging Diverse PreTrained Models from the WildKun Yuan, Hongbo Liu, Mading Li, Muyi Sun 等CVPR 2024 · 被引用 8 次
- QPT-V2: Masked Image Modeling Advances Visual ScoringQizhi Xie, Kun Yuan, Yunpeng Qu, Mingda Wu 等ACM MM 2024 · 被引用 2 次
它引用的顶会 Paper15
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li 等ICCV 2021 · 被引用 1,611 次
相关 Paper
- A Deep Learning based No-reference Quality Assessment Model for UGC VideosWei Sun, Xiongkuo Min, Wei Lu, Guangtao ZhaiACM MM 2022 · 被引用 239 次
- Patch-VQ: 'Patching Up' the Video Quality ProblemZhenqiang Ying, Maniratnam Mandal, Deepti Ghadiyaram, Alan C. BovikCVPR 2021
- Long Short-term Convolutional Transformer for No-Reference Video Quality AssessmentJunyong YouACM MM 2021 · 被引用 46 次
- Perceptual Quality Assessment of Internet VideosJiahua Xu, Jing Li, Xingguang Zhou, Wei Zhou 等ACM MM 2021 · 被引用 44 次
- Blind Natural Video Quality Prediction via Statistical Temporal Features and Deep Spatial FeaturesJari Korhonen, Yicheng Su, Junyong YouACM MM 2020 · 被引用 88 次
