Collaborative Noisy Label Cleaner: Learning Scene-aware Trailers for Multi-modal Highlight Detection in Movies
Bei Gan, Xiujun Shu, Ruizhi Qiao, Haoqian Wu, Keyu Chen, Hanjun Li, Bo Ren
Abstract
Movie highlights stand out of the screenplay for efficient browsing and play a crucial role on social media platforms. Based on existing efforts, this work has two observations: (1) For different annotators, labeling highlight has uncertainty, which leads to inaccurate and time-consuming annotations. (2) Besides previous supervised or unsupervised settings, some existing video corpora can be useful, e.g., trailers, but they are often noisy and incomplete to cover the full highlights. In this work, we study a more practical and promising setting, i.e., reformulating highlight detection as "learning with noisy labels". This setting does not require time-consuming manual annotations and can fully utilize existing abundant video corpora. First, based on movie trailers, we leverage scene segmentation to obtain complete shots, which are regarded as noisy labels. Then, we propose a Collaborative noisy Label Cleaner (CLC) framework to learn from noisy highlight moments. CLC consists of two modules: augmented cross-propagation (ACP) and multi-modality cleaning (MMC). The former aims to exploit the closely related audio-visual signals and fuse them to learn unified multi-modal representations. The latter aims to achieve cleaner highlight labels by observing the changes in losses among different modalities. To verify the effectiveness of CLC, we further collect a large-scale highlight dataset named MovieLights. Comprehensive experiments on MovieLights and YouTube Highlights datasets demonstrate the effectiveness of our approach. Code has been made available at: https://github.com/TencentYoutuResearch/HighlightDetection-CLC
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9b81e39-e963-4851-84ed-cba781b8ac7fCited by top-tier papers6
- Visual Objectification in Films: Towards a New AI Task for Video InterpretationJulie Tores, Lucile Sassatelli, Hui-Yin Wu, Clement Bergman et al.CVPR 2024 · 3 citations
- An Inverse Partial Optimal Transport Framework for Music-guided Trailer GenerationYutong Wang, Sidan Zhu, Hongteng Xu, Dixin LuoACM MM 2024 · 2 citations
- TVHighlights: LLM-Guided Human-Free Collaborative Training for Video Highlight Detection in Movies and TV DramasQi Qiu, Xuan Wu, Jiawei Peng, Yuan Miao et al.CVPR 2026
- VlogReward: Learning Multi-Dimensional Evaluation for Vlog EditingYexiang Liu, Wen Zhong, Sijie Zhu, Xin Gu et al.ICML 2026
- Video Scene Segmentation with Genre and Duration SignalsJungu Cho, Seong Jong Ha, Hae-Gon JeonICLR 2026
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 1,326 citations
- Symmetric Cross Entropy for Robust Learning With Noisy LabelsYisen Wang, Xingjun Ma, Zaiyi Chen, Yuan Luo et al.ICCV 2019 · 1,125 citations
- Normalized Loss Functions for Deep Learning with Noisy LabelsXingjun Ma, Hanxun Huang, Yisen Wang, Simone Romano et al.ICML 2020 · 547 citations
Related papers
- UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight DetectionYe Liu, Siyuan Li, Yang Wu, Chang Wen Chen et al.CVPR 2022 · 150 citations
- Cross-Category Highlight Detection via Feature Decomposition and Modality AlignmentZhenduo ZhangAAAI 2023 · 2 citations
- Cross-category Video Highlight Detection via Set-based LearningMinghao Xu, Hang Wang, Bingbing Ni, Riheng Zhu et al.ICCV 2021 · 63 citations
- Temporal Cue Guided Video Highlight Detection with Low-Rank Audio-Visual FusionQinghao Ye, Xiyue Shen, Yuan Gao, Zirui Wang et al.ICCV 2021 · 58 citations
- MS-DETR: Towards Effective Video Moment Retrieval and Highlight Detection by Joint Motion-Semantic LearningHongxu Ma, Guanshuo Wang, Fufu Yu, Qiong Jia et al.ACM MM 2025 · 9 citations
