TASED-Net: Temporally-Aggregating Spatial Encoder-Decoder Network for Video Saliency Detection
Kyle Min, Jason J. Corso
Abstract
TASED-Net is a 3D fully-convolutional network architecture for video saliency detection. It consists of two building blocks: first, the encoder network extracts low-resolution spatiotemporal features from an input clip of several consecutive frames, and then the following prediction network decodes the encoded features spatially while aggregating all the temporal information. As a result, a single prediction map is produced from an input clip of multiple frames. Frame-wise saliency maps can be predicted by applying TASED-Net in a sliding-window fashion to a video. The proposed approach assumes that the saliency map of any frame can be predicted by considering a limited number of past frames. The results of our extensive experiments on video saliency detection validate this assumption and demonstrate that our fully-convolutional model with temporal aggregation method is effective. TASED-Net significantly outperforms previous state-of-the-art approaches on all three major large-scale datasets of video saliency detection: DHF1K, Hollywood2, and UCFSports. After analyzing the results qualitatively, we observe that our model is especially better at attending to salient moving objects.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- Two-Stream Network for Sign Language Recognition and TranslationYutong Chen, Ronglai Zuo, Fangyun Wei, Yu Wu et al.NeurIPS 2022 · 288 citations
- Full-Duplex Strategy for Video Object SegmentationGe-Peng Ji, Keren Fu, Zhe Wu, Deng-Ping Fan et al.ICCV 2021 · 173 citations
- MEDIRL: Predicting the Visual Attention of Drivers via Maximum Entropy Deep Inverse Reinforcement LearningSonia Baee, Erfan Pakdamanian, Inki Kim, Lu Feng et al.ICCV 2021 · 65 citations
- DRIVE: Deep Reinforced Accident Anticipation with Visual ExplanationWentao Bao, Qi Yu, Yu KongICCV 2021 · 64 citations
- Learning Pixel-Level Distinctions for Video Highlight DetectionFanyue Wei, Biao Wang, Tiezheng Ge, Yuning Jiang et al.CVPR 2022 · 26 citations
Related papers
- SalSAC: A Video Saliency Prediction Model with Shuffled Attentions and Correlation-Based ConvLSTMXinyi Wu, Zhenyao Wu, Jinglin Zhang, Lili Ju et al.AAAI 2020 · 73 citations
- Motion Guided Attention for Video Salient Object DetectionHaofeng Li, Guanqi Chen, Guanbin Li, Yizhou YuICCV 2019 · 200 citations
- Dynamic Context-Sensitive Filtering Network for Video Salient Object DetectionMiao Zhang, Jie Liu, Yifei Wang, Yongri Piao et al.ICCV 2021 · 112 citations
- Finding Action Tubes with a Sparse-to-Dense FrameworkYuxi Li, Weiyao Lin, Tao Wang, John See et al.AAAI 2020 · 18 citations
- Hierarchical Self-Attention Network for Action Localization in VideosRizard Renanda Adhi Pramono, Yie-Tarng Chen, Wen-Hsien FangICCV 2019 · 41 citations
