Semantic-Aware and Quality-Aware Interaction Network for Blind Video Quality Assessment
Jianjun Xiang, Yuanjie Dang, Peng Chen, Ronghua Liang, Ruohong Huan, Nan Gao
Abstract
Current state-of-the-art video quality assessment (VQA) models typically integrate various perceptual features to comprehensively represent video quality degradation. These models either directly concatenate features or fuse different perceptual scores while ignoring the domain gaps between cross-aware features, thus failing to adequately learn the correlations and interactions between different perceptual features. To this end, we analyze the independent effects and information gaps of quality-and semantic-aware features on video quality. Based on an analysis of the spatial and temporal differences between two aware features, we propose a semantic-Aware and quality-Aware Interaction Network (A2INet) for blind VQA. For spatial gaps, we introduce a cross-aware guided interaction module to enhance the interaction between semantic-and quality-aware features in a local-to-global manner. Considering temporal discrepancies, we design a cross-aware temporal modeling module to further perceive temporal content variation and quality saliency information, and perceptual features are regressed into quality score by a temporal network and a temporal pooling. Extensive experiments on six benchmark VQA datasets show that our model achieves state-of-the-art performance, and ablation studies further validate the effectiveness of each module. We also present a simple video sampling strategy to balance the effectiveness and efficiency of the model. The code for the proposed method will be released at https://github.com/JianjunXiang/A2INet.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 52dd8af7-2ab1-46c5-95d6-0be05d1391f1Cited by top-tier papers2
- MVQA-68K: A Multi-dimensional and Causally-annotated Dataset with Quality Interpretability for Video AssessmentYanyun Pu, Kehan Li, Zeyi Huang, Zhijie Zhong et al.ACM MM 2025 · 2 citations
- DSP-PCQA: Integrating Multiple Perception Preferences for Point Cloud Quality AssessmentMingxuan Li, Fazhan Zhang, Zhenzhe Hou, Zihao Huang et al.AAAI 2026
Related papers
- Modular Blind Video Quality AssessmentWen Wen, Mu Li, Yabin Zhang, Yiting Liao et al.CVPR 2024 · 23 citations
- A Deep Learning based No-reference Quality Assessment Model for UGC VideosWei Sun, Xiongkuo Min, Wei Lu, Guangtao ZhaiACM MM 2022 · 239 citations
- Blind Natural Video Quality Prediction via Statistical Temporal Features and Deep Spatial FeaturesJari Korhonen, Yicheng Su, Junyong YouACM MM 2020 · 88 citations
- ADGNet: Attention Discrepancy Guided Deep Neural Network for Blind Image Quality AssessmentXiaoyu Ma, Yaqi Wang, Chang Liu, Suiyu Zhang et al.ACM MM 2022 · 6 citations
- Multiview Contrastive Learning for Completely Blind Video Quality Assessment of User Generated ContentShankhanil Mitra, Rajiv SoundararajanACM MM 2022 · 9 citations
