CASP: Consistency-aware Audio-induced Saliency Prediction Model for Omnidirectional Video
Zhaolin Wan, Han Qin, Zhiyang Li, Xiaopeng Fan, Wangmeng Zuo, Debin Zhao
Abstract
Omnidirectional videos (ODVs) present distinct challenges for accurate audio-visual saliency prediction due to their immersive nature, which combines spatial audio with panoramic visuals to enhance the user experience. While auditory cues are crucial for guiding visual attention across the panoramic scene, the interaction between audio and visual stimuli in ODVs remains underexplored. Existing models primarily focus on spatiotemporal visual cues and treat audio signals separately from their spatial and temporal contexts, often leading to misalignments between audio and visual content and undermining temporal consistency across frames. To bridge these gaps, we propose a novel audio-induced saliency prediction model for ODVs that holistically integrates audio and visual inputs through a multi-modal encoder, an audio-visual interaction module, and an audio-visual transformer. Unlike conventional methods that isolate audio cue locations and attributes, our model employs a query-based framework, where learnable audio queries capture comprehensive audio-visual dependencies, thus enhancing saliency prediction by dynamically aligning with audio cues. Besides, we introduce a novel consistency loss to enforce temporal coherence in saliency regions across frames. Extensive experiments demonstrate that our model outperforms state-of-the-art methods in predicting audio-visual salient regions in ODVs, establishing its robustness and superior performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 22eb959e-e3bb-47ce-a5a4-2f8a846474f3Builds on7
- TASED-Net: Temporally-Aggregating Spatial Encoder-Decoder Network for Video Saliency DetectionKyle Min, Jason J. CorsoICCV 2019 · 189 citations
- SalSAC: A Video Saliency Prediction Model with Shuffled Attentions and Correlation-Based ConvLSTMXinyi Wu, Zhenyao Wu, Jinglin Zhang, Lili Ju et al.AAAI 2020 · 73 citations
- DiffSal: Joint Audio and Video Learning for Diffusion Saliency PredictionJunwen Xiong, Peng Zhang, Tao You, Chuanyue Li et al.CVPR 2024 · 12 citations
- Introducing 3D Thumbnails to Access 360-Degree Videos in Virtual RealityAlissa Vermast, Wolfgang HürstIEEE VR 2023 · 10 citations
- Audio-Visual Spatial Integration and Recursive Attention for Robust Sound Source LocalizationSung Jin Um, Dongjin Kim, Jung Uk KimACM MM 2023 · 4 citations
Related papers
- CASP-Net: Rethinking Video Saliency Prediction from an Audio-Visual Consistency Perceptual PerspectiveJunwen Xiong, Ganglai Wang, Peng Zhang, Wei Huang et al.CVPR 2023
- Instance-Level Panoramic Audio-Visual Saliency Detection and RankingRuohao Guo, Dantong Niu, Liao Qu, Yanyu Qi et al.ACM MM 2024
- Panoramic Video Salient Object Detection with Ambisonic Audio GuidanceXiang Li, Haoyuan Cao, Shijie Zhao, Junlin Li et al.AAAI 2023 · 14 citations
- Exploiting Audio-Visual Consistency with Partial Supervision for Spatial Audio GenerationYan-Bo Lin, Yu-Chiang Frank WangAAAI 2021 · 25 citations
- Boosting Audio Visual Question Answering via Key Semantic-Aware CuesGuangyao Li, Henghui Du, Di HuACM MM 2024 · 16 citations
