Learning Pixel-Level Distinctions for Video Highlight Detection
Fanyue Wei, Biao Wang, Tiezheng Ge, Yuning Jiang, Wen Li, Lixin Duan
Abstract
The goal of video highlight detection is to select the most attractive segments from a long video to depict the most interesting parts of the video. Existing methods typically focus on modeling relationship between different video segments in order to learning a model that can assign highlight scores to these segments; however, these approaches do not explicitly consider the contextual dependency within individual segments. To this end, we propose to learn pixel-level distinctions to improve the video highlight detection. This pixel-level distinction indicates whether or not each pixel in one video belongs to an interesting section. The advantages of modeling such fine-level distinctions are two-fold. First, it allows us to exploit the temporal and spatial relations of the content in one video, since the distinction of a pixel in one frame is highly dependent on both the content before this frame and the content around this pixel in this frame. Second, learning the pixel-level distinction also gives a good explanation to the video highlight task regarding what contents in a highlight segment will be attractive to people. We design an encoder-decoder network to estimate the pixel-level distinction, in which we leverage the 3D convolutional neural networks to exploit the temporal context information, and further take advantage of the visual saliency to model the spatial distinction. State-of-the-art performance on three public benchmarks clearly validates the effectiveness of our framework for video highlight detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Task-Driven Exploration: Decoupling and Inter-Task Feedback for Joint Moment Retrieval and Highlight DetectionJin Yang, Ping Wei, Huan Li, Ziyang RenCVPR 2024 · 13 citations
- Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight DetectionSung Jin Um, Dongjin Kim, Sangmin Lee, Jung Uk KimAAAI 2025 · 8 citations
- Short Video Segment-level User Dynamic Interests Modeling in Personalized RecommendationZhiyu He, Zhixin Ling, Jiayu Li, Zhiqiang Guo et al.SIGIR 2025 · 4 citations
- CVA: Context-aware Video-text Alignment for Video Temporal GroundingSungho Moon, Seunghun Lee, Jiwan Seo, Sunghoon ImCVPR 2026 · 4 citations
- Cross-Category Highlight Detection via Feature Decomposition and Modality AlignmentZhenduo ZhangAAAI 2023 · 2 citations
Builds on5
- TASED-Net: Temporally-Aggregating Spatial Encoder-Decoder Network for Video Saliency DetectionKyle Min, Jason J. CorsoICCV 2019 · 189 citations
- Joint Visual and Audio Learning for Video Highlight DetectionTaivanbat Badamdorj, Mrigank Rochan, Yang Wang, Li ChengICCV 2021 · 91 citations
- Cross-category Video Highlight Detection via Set-based LearningMinghao Xu, Hang Wang, Bingbing Ni, Riheng Zhu et al.ICCV 2021 · 63 citations
- Find Objects and Focus on Highlights: Mining Object Semantics for Video Highlight Detection via Graph Neural NetworksYingying Zhang, Junyu Gao, Xiaoshan Yang, Chang Liu et al.AAAI 2020 · 15 citations
- STAViS: Spatio-Temporal AudioVisual Saliency NetworkAntigoni Tsiami, Petros Koutras, Petros MaragosCVPR 2020
Related papers
- MS-DETR: Towards Effective Video Moment Retrieval and Highlight Detection by Joint Motion-Semantic LearningHongxu Ma, Guanshuo Wang, Fufu Yu, Qiong Jia et al.ACM MM 2025 · 9 citations
- Temporal Cue Guided Video Highlight Detection with Low-Rank Audio-Visual FusionQinghao Ye, Xiyue Shen, Yuan Gao, Zirui Wang et al.ICCV 2021 · 58 citations
- PR-Net: Preference Reasoning for Personalized Video Highlight DetectionRunnan Chen, Penghao Zhou, Wenzhe Wang, Nenglun Chen et al.ICCV 2021 · 14 citations
- Contrastive Learning for Unsupervised Video Highlight DetectionTaivanbat Badamdorj, Mrigank Rochan, Yang Wang, Li ChengCVPR 2022 · 39 citations
- Graph Attention Based Proposal 3D ConvNets for Action DetectionJin Li, Xianglong Liu, Zhuofan Zong, Wanru Zhao et al.AAAI 2020 · 59 citations
