Efficient Movie Scene Detection using State-Space Transformers
Md Mohaiminul Islam, Mahmudul Hasan, Kishan Shamsundar Athrey, Tony Braskich, Gedas Bertasius
Abstract
The ability to distinguish between different movie scenes is critical for understanding the storyline of a movie. However, accurately detecting movie scenes is often challenging as it requires the ability to reason over very long movie segments. This contrasts with most existing video recognition models, which are typically designed for short-range video analysis. This work proposes a State-Space Transformer model that can efficiently capture dependencies in long movie videos for accurate movie scene detection. Our model, called TranS4mer, is built using a novel S4A building block, combining the strengths of structured state-space sequence (S4) and self-attention (A) layers. Given a sequence of frames divided into movie shots (uninterrupted periods where the camera position does not change), the S4A block first applies self-attention to capture short-range intra-shot dependencies. Afterward, the state-space operation in the S4A block aggregates long-range inter-shot cues. The final TranS4mer model, which can be trained end-to-end, is obtained by stacking the S4A blocks one after the other multiple times. Our proposed TranS4mer outperforms all prior methods in three movie scene detection datasets, including MovieNet, BBC, and OVSD, while being 2× faster and requiring 3× less GPU memory than standard Transformer models. We will release our code and models. * Research done while MI was an intern at Comcast Labs. movie scenes is essential for understanding the plot of a movie. Moreover, identifying movie scenes enables broader applications, such as content-driven video search, preview generation, and minimally disruptive ad insertion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers27
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
- FlashFFTConv: Efficient Convolutions for Long Sequences with Tensor CoresDaniel Y. Fu, Hermann Kumbong, Eric Nguyen, Christopher RéICLR 2024 · 41 citations
- A Simple LLM Framework for Long-Range Video Question-AnsweringCe Zhang, Taixi Lu, Md Mohaiminul Islam, Ziyang Wang et al.EMNLP 2024 · 37 citations
- MambaTrack: A Simple Baseline for Multiple Object Tracking with State Space ModelChangcheng Xiao, Qiong Cao, Zhigang Luo, Long LanACM MM 2024 · 31 citations
- MambaPro: Multi-Modal Object Re-identification with Mamba Aggregation and Synergistic PromptYuhao Wang, Xuehu Liu, Tianyu Yan, Yang Liu et al.AAAI 2025 · 30 citations
Builds on17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
Related papers
- Towards Global Video Scene Segmentation with Context-Aware TransformerYang Yang, Yurui Huang, Weili Guo, Baohua Xu et al.AAAI 2023 · 34 citations
- Frequency-Aware Spatiotemporal Transformers for Video Inpainting DetectionBingyao Yu, Wanhua Li, Xiu Li, Jiwen Lu et al.ICCV 2021 · 38 citations
- Multimodal High-order Relation Transformer for Scene Boundary DetectionXi Wei, Zhangxiang Shi, Tianzhu Zhang, Xiaoyuan Yu et al.ICCV 2023 · 7 citations
- TxVAD: Improved Video Action Detection by TransformersZhenyu Wu, Zhou Ren, Yi Wu, Zhangyang Wang et al.ACM MM 2022 · 5 citations
- MS-TCT: Multi-Scale Temporal ConvTransformer for Action DetectionRui Dai, Srijan Das, Kumara Kahatapitiya, Michael S. Ryoo et al.CVPR 2022 · 93 citations
