OS-MSL: One Stage Multimodal Sequential Link Framework for Scene Segmentation and Classification
Ye Liu, Lingfeng Qiao, Di Yin, Zhuoxuan Jiang, Xinghua Jiang, Deqiang Jiang, Bo Ren
Abstract
Scene segmentation and classification (SSC) serve as a critical step towards the field of video structuring analysis. Intuitively, jointly learning of these two tasks can promote each other by sharing common information. However, scene segmentation concerns more on the local difference between adjacent shots while classification needs the global representation of scene segments, which probably leads to the model dominated by one of the two tasks in the training phase. In this paper, from an alternate perspective to overcome the above challenges, we unite these two tasks into one task by a new form of predicting shots link: a link connects two adjacent shots, indicating that they belong to the same scene or category. To the end, we propose a general One Stage Multimodal Sequential Link Framework (OS-MSL) to both distinguish and leverage the two-fold semantics by reforming the two learning tasks into a unified one. Furthermore, we tailor a specific module called DiffCorrNet to explicitly extract the information of differences and correlations among shots. Extensive experiments on a brand-new large scale dataset collected from real-world applications, and MovieScenes are conducted. Both the results demonstrate the effectiveness of our proposed method against strong baselines. The code is made available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f6fc4591-3eb3-4a7a-959b-1fc2d99e2055Cited by top-tier papers2
- OSAN: A One-Stage Alignment Network to Unify Multimodal Alignment and Unsupervised Domain AdaptationYe Liu, Lingfeng Qiao, Changchong Lu, Di Yin et al.CVPR 2023
- NewsNet: A Novel Dataset for Hierarchical Temporal SegmentationHaoqian Wu, Keyu Chen, Haozhe Liu, Mingchen Zhuge et al.CVPR 2023
Builds on5
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Learnable Optimal Sequential Grouping for Video Scene DetectionDaniel Rotman, Yevgeny Yaroker, Elad Amrani, Udi Barzelay et al.ACM MM 2020 · 12 citations
- A Local-to-Global Approach to Multi-Modal Movie Scene SegmentationAnyi Rao, Linning Xu, Yu Xiong, Guodong Xu et al.CVPR 2020
- Shot Contrastive Self-Supervised Learning for Scene Boundary DetectionShixing Chen, Xiaohan Nie, David Fan, Dongqing Zhang et al.CVPR 2021
Related papers
- Modality-Aware Shot Relating and Comparing for Video Scene DetectionJiawei Tan, Hongxing Wang, Kang Dang, Jiaxin Li et al.AAAI 2025 · 1 citation
- Multimodal High-order Relation Transformer for Scene Boundary DetectionXi Wei, Zhangxiang Shi, Tianzhu Zhang, Xiaoyuan Yu et al.ICCV 2023 · 7 citations
- Scene Consistency Representation Learning for Video Scene SegmentationHaoqian Wu, Keyu Chen, Yanan Luo, Ruizhi Qiao et al.CVPR 2022 · 19 citations
- Neighbor Relations Matter in Video Scene DetectionJiawei Tan, Hongxing Wang, Jiaxin Li, Zhilong Ou et al.CVPR 2024 · 1 citation
- Towards Global Video Scene Segmentation with Context-Aware TransformerYang Yang, Yurui Huang, Weili Guo, Baohua Xu et al.AAAI 2023 · 34 citations
