Cross-Category Highlight Detection via Feature Decomposition and Modality Alignment
Zhenduo Zhang
摘要
Learning an autonomous highlight video detector with good transferability across video categories, called Cross-Category Video Highlight Detection(CC-VHD), is crucial for the practical application on video-based media platforms. To tackle this problem, we first propose a framework that treats the CC-VHD as learning category-independent highlight feature representation. Under this framework, we propose a novel module, named Multi-task Feature Decomposition Branch which jointly conducts label prediction, cyclic feature reconstruction, and adversarial feature reconstruction to decompose the video features into two independent components: highlight-related component and category-related component. Besides, we propose to align the visual and audio modalities to one aligned feature space before conducting modality fusion, which has not been considered in previous works. Finally, the extensive experimental results on three challenging public benchmarks validate the efficacy of our paradigm and the superiority over the existing state-of-the-art approaches to video highlight detection.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- An Empirical Study of Training End-to-End Vision-and-Language TransformersZi-Yi Dou, Yichong Xu, Zhe Gan, Jianfeng Wang 等CVPR 2022 · 被引用 313 次
- Joint Visual and Audio Learning for Video Highlight DetectionTaivanbat Badamdorj, Mrigank Rochan, Yang Wang, Li ChengICCV 2021 · 被引用 91 次
- Cross-category Video Highlight Detection via Set-based LearningMinghao Xu, Hang Wang, Bingbing Ni, Riheng Zhu 等ICCV 2021 · 被引用 63 次
- Disentangled Representation for Age-Invariant Face Recognition: A Mutual Information Minimization PerspectiveXuege Hou, Yali Li, Shengjin WangICCV 2021 · 被引用 31 次
相关 Paper
- Temporal Cue Guided Video Highlight Detection with Low-Rank Audio-Visual FusionQinghao Ye, Xiyue Shen, Yuan Gao, Zirui Wang 等ICCV 2021 · 被引用 58 次
- UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight DetectionYe Liu, Siyuan Li, Yang Wu, Chang Wen Chen 等CVPR 2022 · 被引用 150 次
- Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight DetectionYicheng Xiao, Zhuoyan Luo, Yong Liu, Yue Ma 等CVPR 2024 · 被引用 43 次
- TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight DetectionHao Sun, Mingyao Zhou, Wenjing Chen, Wei XieAAAI 2024
- CMHKF: Cross-Modality Heterogeneous Knowledge Fusion for Weakly Supervised Video Anomaly DetectionGuohua Wang, Shengping Song, Wuchun He, Yongsen ZhengACL 2025 · 被引用 2 次
