Temporal Cue Guided Video Highlight Detection with Low-Rank Audio-Visual Fusion
Qinghao Ye, Xiyue Shen, Yuan Gao, Zirui Wang, Qi Bi, Ping Li, Guang Yang
摘要
Video highlight detection plays an increasingly important role in social media content filtering, however, it remains highly challenging to develop automated video highlight detection methods because of the lack of temporal annotations (i.e., where the highlight moments are in long videos) for supervised learning. In this paper, we propose a novel weakly supervised method that can learn to detect highlights by mining video characteristics with video level annotations (topic tags) only. Particularly, we exploit audio-visual features to enhance video representation and take temporal cues into account for improving detection performance. Our contributions are threefold: 1) we propose an audio-visual tensor fusion mechanism that efficiently models the complex association between two modalities while reducing the gap of the heterogeneity between the two modalities; 2) we introduce a novel hierarchical temporal context encoder to embed local temporal clues in between neighboring segments; 3) finally, we alleviate the gradient vanishing problem theoretically during model optimization with attention-gated instance aggregation. Extensive experiments on two benchmark datasets (YouTube Highlights and TVSum) have demonstrated our method outperforms other state-of-the-art methods with remarkable improvements.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- UniVTG: Towards Unified Video-Language Temporal GroundingKevin Qinghong Lin, Pengchuan Zhang, Joya Chen, Shraman Pramanick 等ICCV 2023 · 被引用 221 次
- UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight DetectionYe Liu, Siyuan Li, Yang Wu, Chang Wen Chen 等CVPR 2022 · 被引用 150 次
- HiTeA: Hierarchical Temporal-Aware Video-Language Pre-trainingQinghao Ye, Guohai Xu, Ming Yan, Haiyang Xu 等ICCV 2023 · 被引用 102 次
- Learning Content-Enhanced Mask Transformer for Domain Generalized Urban-Scene SegmentationQi Bi, Shaodi You, Theo GeversAAAI 2024 · 被引用 77 次
- Learning Frequency-Adapted Vision Foundation Model for Domain Generalized Semantic SegmentationQi Bi, Jingjun Yi, Hao Zheng, Haolan Zhan 等NeurIPS 2024 · 被引用 62 次
相关 Paper
- Joint Visual and Audio Learning for Video Highlight DetectionTaivanbat Badamdorj, Mrigank Rochan, Yang Wang, Li ChengICCV 2021 · 被引用 91 次
- Contrastive Learning for Unsupervised Video Highlight DetectionTaivanbat Badamdorj, Mrigank Rochan, Yang Wang, Li ChengCVPR 2022 · 被引用 39 次
- Cross-Category Highlight Detection via Feature Decomposition and Modality AlignmentZhenduo ZhangAAAI 2023 · 被引用 2 次
- Cross-category Video Highlight Detection via Set-based LearningMinghao Xu, Hang Wang, Bingbing Ni, Riheng Zhu 等ICCV 2021 · 被引用 63 次
- CMHKF: Cross-Modality Heterogeneous Knowledge Fusion for Weakly Supervised Video Anomaly DetectionGuohua Wang, Shengping Song, Wuchun He, Yongsen ZhengACL 2025 · 被引用 2 次
