Joint Visual and Audio Learning for Video Highlight Detection
Taivanbat Badamdorj, Mrigank Rochan, Yang Wang, Li Cheng
Abstract
In video highlight detection, the goal is to identify the interesting moments within an unedited video. Although the audio component of the video provides important cues for highlight detection, the majority of existing efforts focus almost exclusively on the visual component. In this paper, we argue that both audio and visual components of a video should be modeled jointly to retrieve its best moments. To this end, we propose an audio-visual network for video highlight detection. At the core of our approach lies a bimodal attention mechanism, which captures the interaction between the audio and visual components of a video, and produces fused representations to facilitate highlight detection. Furthermore, we introduce a noise sentinel technique to adaptively discount a noisy visual or audio modality. Empirical evaluations on two benchmark datasets demonstrate the superior performance of our approach over the state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b16f97ee-0054-4a4b-92e9-f965dbce2f37Cited by top-tier papers23
- UniVTG: Towards Unified Video-Language Temporal GroundingKevin Qinghong Lin, Pengchuan Zhang, Joya Chen, Shraman Pramanick et al.ICCV 2023 · 221 citations
- UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight DetectionYe Liu, Siyuan Li, Yang Wu, Chang Wen Chen et al.CVPR 2022 · 150 citations
- Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight DetectionYicheng Xiao, Zhuoyan Luo, Yong Liu, Yue Ma et al.CVPR 2024 · 43 citations
- Contrastive Learning for Unsupervised Video Highlight DetectionTaivanbat Badamdorj, Mrigank Rochan, Yang Wang, Li ChengCVPR 2022 · 39 citations
- Hyperbolic Audio-visual Zero-shot LearningJie Hong, Zeeshan Hayder, Junlin Han, Pengfei Fang et al.ICCV 2023 · 27 citations
Builds on2
Related papers
- Temporal Cue Guided Video Highlight Detection with Low-Rank Audio-Visual FusionQinghao Ye, Xiyue Shen, Yuan Gao, Zirui Wang et al.ICCV 2021 · 58 citations
- Cross-Modal Relation-Aware Networks for Audio-Visual Event LocalizationHaoming Xu, Runhao Zeng, Qingyao Wu, Mingkui Tan et al.ACM MM 2020 · 97 citations
- CACE-Net: Co-guidance Attention and Contrastive Enhancement for Effective Audio-Visual Event LocalizationXiang He, Xiangxi Liu, Yang Li, Dongcheng Zhao et al.ACM MM 2024 · 8 citations
- Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight DetectionSung Jin Um, Dongjin Kim, Sangmin Lee, Jung Uk KimAAAI 2025 · 8 citations
- CASP-Net: Rethinking Video Saliency Prediction from an Audio-Visual Consistency Perceptual PerspectiveJunwen Xiong, Ganglai Wang, Peng Zhang, Wei Huang et al.CVPR 2023
