From Semantic Categories to Fixations: A Novel Weakly-Supervised Visual-Auditory Saliency Detection Approach
Guotao Wang, Chenglizhao Chen, Deng-Ping Fan, Aimin Hao, Hong Qin
Abstract
Thanks to the rapid advances in the deep learning techniques and the wide availability of large-scale training sets, the performances of video saliency detection models have been improving steadily and significantly. However, the deep learning based visual-audio fixation prediction is still in its infancy. At present, only a few visual-audio sequences have been furnished with real fixations being recorded in the real visual-audio environment. Hence, it would be neither efficiency nor necessary to re-collect real fixations under the same visual-audio circumstance. To address the problem, this paper advocate a novel approach in a weaklysupervised manner to alleviating the demand of large-scale training sets for visual-audio model training. By using the video category tags only, we propose the selective class activation mapping (SCAM), which follows a coarse-to-fine strategy to select the most discriminative regions in the spatial-temporal-audio circumstance. Moreover, these regions exhibit high consistency with the real human-eye fixations, which could subsequently be employed as the pseudo GTs to train a new spatial-temporal-audio (STA) network. Without resorting to any real fixation, the performance of our STA network is comparable to that of the fully supervised ones. Our code and results are publicly available at https://github.com/guotaowang/STANet .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Synthetic Data Supervised Salient Object DetectionZhenyu Wu, Lin Wang, Wei Wang, Tengfei Shi et al.ACM MM 2022 · 29 citations
- Unsupervised Self-Driving Attention Prediction via Uncertainty Mining and Knowledge EmbeddingPengfei Zhu, Mengshi Qi, Xia Li, Weijian Li et al.ICCV 2023 · 23 citations
- NPF-200: A Multi-Modal Eye Fixation Dataset and Method for Non-Photorealistic VideosZiyu Yang, Sucheng Ren, Zongwei Wu, Nanxuan Zhao et al.ACM MM 2023 · 2 citations
Builds on11
- Reliability Does Matter: An End-to-End Weakly Supervised Semantic Segmentation ApproachBingfeng Zhang, Jimin Xiao, Yunchao Wei, Mingjie Sun et al.AAAI 2020 · 227 citations
- Self-Training and Adversarial Background Regularization for Unsupervised Domain Adaptive One-Stage Object DetectionSeunghyeon Kim, Jaehoon Choi, Taekyung Kim, Changick KimICCV 2019 · 211 citations
- Motion Guided Attention for Video Salient Object DetectionHaofeng Li, Guanqi Chen, Guanbin Li, Yizhou YuICCV 2019 · 200 citations
- TASED-Net: Temporally-Aggregating Spatial Encoder-Decoder Network for Video Saliency DetectionKyle Min, Jason J. CorsoICCV 2019 · 189 citations
- Pyramid Constrained Self-Attention Network for Fast Video Salient Object DetectionYuchao Gu, Lijuan Wang, Ziqin Wang, Yun Liu et al.AAAI 2020 · 184 citations
Related papers
- Keep CALM and Improve Visual Feature AttributionJae-Myung Kim, Junsuk Choe, Zeynep Akata, Seong Joon OhICCV 2021 · 22 citations
- CASP-Net: Rethinking Video Saliency Prediction from an Audio-Visual Consistency Perceptual PerspectiveJunwen Xiong, Ganglai Wang, Peng Zhang, Wei Huang et al.CVPR 2023
- Weakly Supervised Video Salient Object DetectionWangbo Zhao, Jing Zhang, Long Li, Nick Barnes et al.CVPR 2021
- Embedded Discriminative Attention Mechanism for Weakly Supervised Semantic SegmentationTong Wu, Junshi Huang, Guangyu Gao, Xiaoming Wei et al.CVPR 2021
- Boat in the Sky: Background Decoupling and Object-aware Pooling for Weakly Supervised Semantic SegmentationJianjun Xu, Hongtao Xie, Hai Xu, Yuxin Wang et al.ACM MM 2022 · 13 citations
