CASP-Net: Rethinking Video Saliency Prediction from an Audio-Visual Consistency Perceptual Perspective
Junwen Xiong, Ganglai Wang, Peng Zhang, Wei Huang, Yufei Zha, Guangtao Zhai
Abstract
Incorporating the audio stream enables Video Saliency Prediction (VSP) to imitate the selective attention mechanism of human brain. By focusing on the benefits of joint auditory and visual information, most VSP methods are capable of exploiting semantic correlation between vision and audio modalities but ignoring the negative effects due to the temporal inconsistency of audio-visual intrinsics. Inspired by the biological inconsistency-correction within multi-sensory information, in this study, a consistencyaware audio-visual saliency prediction network (CASP-Net) is proposed, which takes a comprehensive consideration of the audio-visual semantic interaction and consistent perception. In addition a two-stream encoder for elegant association between video frames and corresponding sound source, a novel consistency-aware predictive coding is also designed to improve the consistency within audio and visual representations iteratively. To further aggregate the multi-scale audio-visual information, a saliency decoder is introduced for the final saliency map generation. Substantial experiments demonstrate that the proposed CASP-Net outperforms the other state-of-the-art methods on six challenging audio-visual eye-tracking datasets. For a demo of our system please see our project webpage.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9528c26a-ec9c-4f0d-8625-740b9b067201Cited by top-tier papers3
- DiffSal: Joint Audio and Video Learning for Diffusion Saliency PredictionJunwen Xiong, Peng Zhang, Tao You, Chuanyue Li et al.CVPR 2024 · 12 citations
- Attend to Anything: Foundation Model for Unified Human Attention ModelingWenzhuo Zhao, Ronghao Xian, Keren Fu, Qijun ZhaoICML 2026
- CASP: Consistency-aware Audio-induced Saliency Prediction Model for Omnidirectional VideoZhaolin Wan, Han Qin, Zhiyang Li, Xiaopeng Fan et al.CVPR 2025
Builds on5
- TASED-Net: Temporally-Aggregating Spatial Encoder-Decoder Network for Video Saliency DetectionKyle Min, Jason J. CorsoICCV 2019 · 189 citations
- Cross-Modal Attention Network for Temporal Inconsistent Audio-Visual Event LocalizationHanyu Xuan, Zhenyu Zhang, Shuo Chen, Jian Yang et al.AAAI 2020 · 110 citations
- STAViS: Spatio-Temporal AudioVisual Saliency NetworkAntigoni Tsiami, Petros Koutras, Petros MaragosCVPR 2020
- VisualVoice: Audio-Visual Speech Separation With Cross-Modal ConsistencyRuohan Gao, Kristen GraumanCVPR 2021
- Semantic Audio-Visual NavigationChangan Chen, Ziad Al-Halah, Kristen GraumanCVPR 2021
Related papers
- CACE-Net: Co-guidance Attention and Contrastive Enhancement for Effective Audio-Visual Event LocalizationXiang He, Xiangxi Liu, Yang Li, Dongcheng Zhao et al.ACM MM 2024 · 8 citations
- Joint Visual and Audio Learning for Video Highlight DetectionTaivanbat Badamdorj, Mrigank Rochan, Yang Wang, Li ChengICCV 2021 · 91 citations
- SelM: Selective Mechanism based Audio-Visual SegmentationJiaxu Li, Songsong Yu, Yifan Wang, Lijun Wang et al.ACM MM 2024 · 5 citations
- Multi-modal Deepfake Detection via Multi-task Audio-Visual Prompt LearningHui Miao, Yuanfang Guo, Zeming Liu, Yunhong WangAAAI 2025 · 8 citations
- Progressive Spatio-temporal Perception for Audio-Visual Question AnsweringGuangyao Li, Wenxuan Hou, Di HuACM MM 2023 · 39 citations
