SalSAC: A Video Saliency Prediction Model with Shuffled Attentions and Correlation-Based ConvLSTM
Xinyi Wu, Zhenyao Wu, Jinglin Zhang, Lili Ju, Song Wang
摘要
The performance of predicting human fixations in videos has been much enhanced with the help of development of the convolutional neural networks (CNN). In this paper, we propose a novel end-to-end neural network “SalSAC” for video saliency prediction, which uses the CNN-LSTM-Attention as the basic architecture and utilizes the information from both static and dynamic aspects. To better represent the static information of each frame, we first extract multi-level features of same size from different layers of the encoder CNN and calculate the corresponding multi-level attentions, then we randomly shuffle these attention maps among levels and multiply them to the extracted multi-level features respectively. Through this way, we leverage the attention consistency across different layers to improve the robustness of the network. On the dynamic aspect, we propose a correlation-based ConvLSTM to appropriately balance the influence of the current and preceding frames to the prediction. Experimental results on the DHF1K, Hollywood2 and UCF-sports datasets show that SalSAC outperforms many existing state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- NPF-200: A Multi-Modal Eye Fixation Dataset and Method for Non-Photorealistic VideosZiyu Yang, Sucheng Ren, Zongwei Wu, Nanxuan Zhao 等ACM MM 2023 · 被引用 2 次
- CASP: Consistency-aware Audio-induced Saliency Prediction Model for Omnidirectional VideoZhaolin Wan, Han Qin, Zhiyang Li, Xiaopeng Fan 等CVPR 2025
- From Semantic Categories to Fixations: A Novel Weakly-Supervised Visual-Auditory Saliency Detection ApproachGuotao Wang, Chenglizhao Chen, Deng-Ping Fan, Aimin Hao 等CVPR 2021
相关 Paper
- TASED-Net: Temporally-Aggregating Spatial Encoder-Decoder Network for Video Saliency DetectionKyle Min, Jason J. CorsoICCV 2019 · 被引用 189 次
- Hierarchical Self-Attention Network for Action Localization in VideosRizard Renanda Adhi Pramono, Yie-Tarng Chen, Wen-Hsien FangICCV 2019 · 被引用 41 次
- CASP-Net: Rethinking Video Saliency Prediction from an Audio-Visual Consistency Perceptual PerspectiveJunwen Xiong, Ganglai Wang, Peng Zhang, Wei Huang 等CVPR 2023
- CSTA: CNN-based Spatiotemporal Attention for Video SummarizationJaewon Son, Jaehun Park, Kwangsu KimCVPR 2024 · 被引用 17 次
- Structured Modeling of Joint Deep Feature and Prediction Refinement for Salient Object DetectionYingyue Xu, Dan Xu, Xiaopeng Hong, Wanli Ouyang 等ICCV 2019 · 被引用 43 次
