Scene Consistency Representation Learning for Video Scene Segmentation
Haoqian Wu, Keyu Chen, Yanan Luo, Ruizhi Qiao, Bo Ren, Haozhe Liu, Weicheng Xie, Linlin Shen
Abstract
A long-term video, such as a movie or TV show, is composed of various scenes, each of which represents a series of shots sharing the same semantic story. Spotting the correct scene boundary from the long-term video is a challenging task, since a model must understand the storyline of the video to figure out where a scene starts and ends. To this end, we propose an effective Self-Supervised Learning (SSL) framework to learn better shot representations from unlabeled long-term videos. More specifically, we present an SSL scheme to achieve scene consistency, while exploring considerable data augmentation and shuffling methods to boost the model generalizability. Instead of explicitly learning the scene boundary features as in the previous methods, we introduce a vanilla temporal model with less inductive bias to verify the quality of the shot features. Our method achieves the state-of-the-art performance on the task of Video Scene Segmentation. Additionally, we suggest a more fair and reasonable benchmark to evaluate the performance of Video Scene Segmentation methods. The code is made available. <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> https://github.com/TencentYoutuResearch/SceneSegmentation-SCRL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 187d80ec-8833-495c-9027-8d273c87da2dCited by top-tier papers15
- Towards Global Video Scene Segmentation with Context-Aware TransformerYang Yang, Yurui Huang, Weili Guo, Baohua Xu et al.AAAI 2023 · 34 citations
- MultiShotMaster: A Controllable Multi-Shot Video Generation FrameworkQinghe Wang, Xiaoyu Shi, Baolu Li, Weikang Bian et al.CVPR 2026 · 33 citations
- SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision ViewersZheng Ning, Brianna L. Wimer, Kaiwen Jiang, Keyi Chen et al.CHI 2024 · 27 citations
- Multimodal High-order Relation Transformer for Scene Boundary DetectionXi Wei, Zhangxiang Shi, Tianzhu Zhang, Xiaoyuan Yu et al.ICCV 2023 · 7 citations
- MEGA: Multimodal Alignment Aggregation and Distillation For Cinematic Video SegmentationNajmeh Sadoughi, Xinyu Li, Avijit Vajpayee, David Fan et al.ICCV 2023 · 6 citations
Builds on19
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi et al.NeurIPS 2020 · 2,611 citations
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
Related papers
- Video Scene Segmentation with Genre and Duration SignalsJungu Cho, Seong Jong Ha, Hae-Gon JeonICLR 2026
- Unsupervised Object-Level Representation Learning from Scene ImagesJiahao Xie, Xiaohang Zhan, Ziwei Liu, Yew Soon Ong et al.NeurIPS 2021 · 93 citations
- Spatiotemporal Contrastive Video Representation LearningRui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang et al.CVPR 2021
- Tracklet Self-Supervised Learning for Unsupervised Person Re-IdentificationGuile Wu, Xiatian Zhu, Shaogang GongAAAI 2020 · 97 citations
- OS-MSL: One Stage Multimodal Sequential Link Framework for Scene Segmentation and ClassificationYe Liu, Lingfeng Qiao, Di Yin, Zhuoxuan Jiang et al.ACM MM 2022 · 5 citations
