From Semantic Categories to Fixations: A Novel Weakly-Supervised Visual-Auditory Saliency Detection Approach
Guotao Wang, Chenglizhao Chen, Deng-Ping Fan, Aimin Hao, Hong Qin
摘要
Thanks to the rapid advances in the deep learning techniques and the wide availability of large-scale training sets, the performances of video saliency detection models have been improving steadily and significantly. However, the deep learning based visual-audio fixation prediction is still in its infancy. At present, only a few visual-audio sequences have been furnished with real fixations being recorded in the real visual-audio environment. Hence, it would be neither efficiency nor necessary to re-collect real fixations under the same visual-audio circumstance. To address the problem, this paper advocate a novel approach in a weaklysupervised manner to alleviating the demand of large-scale training sets for visual-audio model training. By using the video category tags only, we propose the selective class activation mapping (SCAM), which follows a coarse-to-fine strategy to select the most discriminative regions in the spatial-temporal-audio circumstance. Moreover, these regions exhibit high consistency with the real human-eye fixations, which could subsequently be employed as the pseudo GTs to train a new spatial-temporal-audio (STA) network. Without resorting to any real fixation, the performance of our STA network is comparable to that of the fully supervised ones. Our code and results are publicly available at https://github.com/guotaowang/STANet .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Synthetic Data Supervised Salient Object DetectionZhenyu Wu, Lin Wang, Wei Wang, Tengfei Shi 等ACM MM 2022 · 被引用 29 次
- Unsupervised Self-Driving Attention Prediction via Uncertainty Mining and Knowledge EmbeddingPengfei Zhu, Mengshi Qi, Xia Li, Weijian Li 等ICCV 2023 · 被引用 23 次
- NPF-200: A Multi-Modal Eye Fixation Dataset and Method for Non-Photorealistic VideosZiyu Yang, Sucheng Ren, Zongwei Wu, Nanxuan Zhao 等ACM MM 2023 · 被引用 2 次
它引用的顶会 Paper11
- Reliability Does Matter: An End-to-End Weakly Supervised Semantic Segmentation ApproachBingfeng Zhang, Jimin Xiao, Yunchao Wei, Mingjie Sun 等AAAI 2020 · 被引用 227 次
- Self-Training and Adversarial Background Regularization for Unsupervised Domain Adaptive One-Stage Object DetectionSeunghyeon Kim, Jaehoon Choi, Taekyung Kim, Changick KimICCV 2019 · 被引用 211 次
- Motion Guided Attention for Video Salient Object DetectionHaofeng Li, Guanqi Chen, Guanbin Li, Yizhou YuICCV 2019 · 被引用 200 次
- TASED-Net: Temporally-Aggregating Spatial Encoder-Decoder Network for Video Saliency DetectionKyle Min, Jason J. CorsoICCV 2019 · 被引用 189 次
- Pyramid Constrained Self-Attention Network for Fast Video Salient Object DetectionYuchao Gu, Lijuan Wang, Ziqin Wang, Yun Liu 等AAAI 2020 · 被引用 184 次
相关 Paper
- Keep CALM and Improve Visual Feature AttributionJae-Myung Kim, Junsuk Choe, Zeynep Akata, Seong Joon OhICCV 2021 · 被引用 22 次
- CASP-Net: Rethinking Video Saliency Prediction from an Audio-Visual Consistency Perceptual PerspectiveJunwen Xiong, Ganglai Wang, Peng Zhang, Wei Huang 等CVPR 2023
- Weakly Supervised Video Salient Object DetectionWangbo Zhao, Jing Zhang, Long Li, Nick Barnes 等CVPR 2021
- Embedded Discriminative Attention Mechanism for Weakly Supervised Semantic SegmentationTong Wu, Junshi Huang, Guangyu Gao, Xiaoming Wei 等CVPR 2021
- Boat in the Sky: Background Decoupling and Object-aware Pooling for Weakly Supervised Semantic SegmentationJianjun Xu, Hongtao Xie, Hai Xu, Yuxin Wang 等ACM MM 2022 · 被引用 13 次
