Contrastive Learning for Unsupervised Video Highlight Detection
Taivanbat Badamdorj, Mrigank Rochan, Yang Wang, Li Cheng
Abstract
Video highlight detection can greatly simplify video browsing, potentially paving the way for a wide range of ap-plications. Existing efforts are mostly fully-supervised, requiring humans to manually identify and label the interesting moments (called highlights) in a video. Recent weakly supervised methods forgo the use of highlight annotations, but typically require extensive efforts in collecting external data such as web-crawled videos for model learning. This observation has inspired us to consider unsupervised highlight detection where neither frame-level nor video-level annotations are available in training. We propose a simple contrastive learning framework for unsupervised highlight detection. Our framework encodes a video into a vector representation by learning to pick video clips that help to distinguish it from other videos via a contrastive objective using dropout noise. This inherently allows our framework to identify video clips corresponding to highlight of the video. Extensive empirical evaluations on three highlight detection benchmarks demonstrate the superior performance of our approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1629e117-52b7-4493-a07a-5544d6401b3bCited by top-tier papers14
- Prior Knowledge Integration via LLM Encoding and Pseudo Event Regulation for Video Moment RetrievalYiyang Jiang, Wengyu Zhang, Xulu Zhang, Xiaoyong Wei et al.ACM MM 2024 · 21 citations
- Task-Driven Exploration: Decoupling and Inter-Task Feedback for Joint Moment Retrieval and Highlight DetectionJin Yang, Ping Wei, Huan Li, Ziyang RenCVPR 2024 · 13 citations
- MS-DETR: Towards Effective Video Moment Retrieval and Highlight Detection by Joint Motion-Semantic LearningHongxu Ma, Guanshuo Wang, Fufu Yu, Qiong Jia et al.ACM MM 2025 · 9 citations
- Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight DetectionSung Jin Um, Dongjin Kim, Sangmin Lee, Jung Uk KimAAAI 2025 · 8 citations
- CVA: Context-aware Video-text Alignment for Video Temporal GroundingSungho Moon, Seunghun Lee, Jiwan Seo, Sunghoon ImCVPR 2026 · 4 citations
Builds on2
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Joint Visual and Audio Learning for Video Highlight DetectionTaivanbat Badamdorj, Mrigank Rochan, Yang Wang, Li ChengICCV 2021 · 91 citations
Related papers
- Weakly Supervised Video Moment Localization with Contrastive Negative Sample MiningMinghang Zheng, Yanjie Huang, Qingchao Chen, Yang LiuAAAI 2022 · 109 citations
- AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video HashingNiu Lian, Jun Li, Jinpeng Wang, Ruisheng Luo et al.CVPR 2025
- Cross-category Video Highlight Detection via Set-based LearningMinghao Xu, Hang Wang, Bingbing Ni, Riheng Zhu et al.ICCV 2021 · 63 citations
- PR-Net: Preference Reasoning for Personalized Video Highlight DetectionRunnan Chen, Penghao Zhou, Wenzhe Wang, Nenglun Chen et al.ICCV 2021 · 14 citations
- Temporal Cue Guided Video Highlight Detection with Low-Rank Audio-Visual FusionQinghao Ye, Xiyue Shen, Yuan Gao, Zirui Wang et al.ICCV 2021 · 58 citations
