Learnable Optimal Sequential Grouping for Video Scene Detection
Daniel Rotman, Yevgeny Yaroker, Elad Amrani, Udi Barzelay, Rami Ben-Ari
Abstract
Video scene detection is the task of dividing videos into temporal semantic chapters. This is an important preliminary step before attempting to analyze heterogeneous video content. Recently, Optimal Sequential Grouping (OSG) was proposed as a powerful unsupervised solution to solve a formulation of the video scene detection problem. In this work, we extend the capabilities of OSG to the learning regime. By giving the capability to both learn from examples and leverage a robust optimization formulation, we can boost performance and enhance the versatility of the technology. We present a comprehensive analysis of incorporating OSG into deep learning neural networks under various configurations. These configurations include learning an embedding in a straight-forward manner, a tailored loss designed to guide the solution of OSG, and an integrated model where the learning is performed through the OSG pipeline. With thorough evaluation and analysis, we assess the benefits and behavior of the various configurations, and show that our learnable OSG approach exhibits desirable behavior and enhanced performance compared to the state of the art.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e8d5b300-9677-4b95-9253-5d27a925f296Cited by top-tier papers1
Ask how each one uses itRelated papers
- DyStaB: Unsupervised Object Segmentation via Dynamic-Static BootstrappingYanchao Yang, Brian Lai, Stefano SoattoCVPR 2021
- SegSort: Segmentation by Discriminative Sorting of SegmentsJyh-Jing Hwang, Stella X. Yu, Jianbo Shi, Maxwell D. Collins et al.ICCV 2019 · 160 citations
- Learning Video Object Segmentation From Unlabeled VideosXiankai Lu, Wenguan Wang, Jianbing Shen, Yu-Wing Tai et al.CVPR 2020
- Unsupervised Learning From Video With Deep Neural EmbeddingsChengxu Zhuang, Tianwei She, Alex Andonian, Max Sobol Mark et al.CVPR 2020
- Unsupervised Temporal Video Grounding with Deep Semantic ClusteringDaizong Liu, Xiaoye Qu, Yinzhen Wang, Xing Di et al.AAAI 2022 · 52 citations
