STAViS: Spatio-Temporal AudioVisual Saliency Network
Antigoni Tsiami, Petros Koutras, Petros Maragos
Abstract
We introduce STAViS 1 , a spatio-temporal audiovisual saliency network that combines spatio-temporal visual and auditory information in order to efficiently address the problem of saliency estimation in videos. Our approach employs a single network that combines visual saliency and auditory features and learns to appropriately localize sound sources and to fuse the two saliencies in order to obtain a final saliency map. The network has been designed, trained end-to-end, and evaluated on six different databases that contain audiovisual eye-tracking data of a large variety of videos. We compare our method against 8 different stateof-the-art visual saliency models. Evaluation results across databases indicate that our STAViS model outperforms our visual only variant as well as the other state-of-the-art models in the majority of cases. Also, the consistently good performance it achieves for all databases indicates that it is appropriate for estimating saliency "in-the-wild". The code is available at https://github.com/atsiami/STAViS .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Learning Pixel-Level Distinctions for Video Highlight DetectionFanyue Wei, Biao Wang, Tiezheng Ge, Yuning Jiang et al.CVPR 2022 · 26 citations
- CaRDiff: Video Salient Object Ranking Chain of Thought Reasoning for Saliency Prediction with DiffusionYunlong Tang, Gen Zhan, Li Yang, Yiting Liao et al.AAAI 2025 · 16 citations
- Panoramic Video Salient Object Detection with Ambisonic Audio GuidanceXiang Li, Haoyuan Cao, Shijie Zhao, Junlin Li et al.AAAI 2023 · 14 citations
- DiffSal: Joint Audio and Video Learning for Diffusion Saliency PredictionJunwen Xiong, Peng Zhang, Tao You, Chuanyue Li et al.CVPR 2024 · 12 citations
- Everything at Once - Multi-modal Fusion Transformer for Video RetrievalNina Shvetsova, Brian Chen, Andrew Rouditchenko, Samuel Thomas et al.CVPR 2022 · 4 citations
Builds on2
Related papers
- CASP-Net: Rethinking Video Saliency Prediction from an Audio-Visual Consistency Perceptual PerspectiveJunwen Xiong, Ganglai Wang, Peng Zhang, Wei Huang et al.CVPR 2023
- Self-Supervised Moving Vehicle Tracking With Stereo SoundChuang Gan, Hang Zhao, Peihao Chen, David D. Cox et al.ICCV 2019 · 157 citations
- Binaural Audio-Visual LocalizationXinyi Wu, Zhenyao Wu, Lili Ju, Song WangAAAI 2021 · 32 citations
- Cross-modal Background Suppression for Audio-Visual Event LocalizationYan Xia, Zhou ZhaoCVPR 2022 · 62 citations
- Joint Visual and Audio Learning for Video Highlight DetectionTaivanbat Badamdorj, Mrigank Rochan, Yang Wang, Li ChengICCV 2021 · 91 citations
