Multimodal High-order Relation Transformer for Scene Boundary Detection
Xi Wei, Zhangxiang Shi, Tianzhu Zhang, Xiaoyuan Yu, Lei Xiao
Abstract
Scene boundary detection breaks down long videos into meaningful story-telling units and plays a crucial role in high-level video understanding. Despite significant advancements in this area, this task remains a challenging problem as it requires a comprehensive understanding of multimodal cues and high-level semantics. To tackle this issue, we propose a multimodal high-order relation transformer, which integrates a high-order encoder and an adaptive decoder in a unified framework. By modeling the mul-timodal cues and exploring similarities between the shots, the encoder is capable of capturing high-order relations between shots and extracting shot features with context semantics. By clustering the shots adaptively, the decoder can discover more universal switch pattern between successive scenes, thus helping scene boundary detection. Extensive experimental results on three standard benchmarks demonstrate that the proposed model performs favorably against state-of-the-art video scene detection methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3620d83b-afc5-4092-ad0a-868edffbaa32Cited by top-tier papers4
- Modality-Aware Shot Relating and Comparing for Video Scene DetectionJiawei Tan, Hongxing Wang, Kang Dang, Jiaxin Li et al.AAAI 2025 · 1 citation
- Neighbor Relations Matter in Video Scene DetectionJiawei Tan, Hongxing Wang, Jiaxin Li, Zhilong Ou et al.CVPR 2024 · 1 citation
- Video Scene Segmentation with Genre and Duration SignalsJungu Cho, Seong Jong Ha, Hae-Gon JeonICLR 2026
- Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMsLucas Ventura, Antoine Yang, Cordelia Schmid, Gül VarolCVPR 2025
Builds on7
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Fast and Unsupervised Action Boundary Detection for Action SegmentationZexing Du, Xue Wang, Guoqing Zhou, Qing WangCVPR 2022 · 39 citations
- UBoCo: Unsupervised Boundary Contrastive Learning for Generic Event Boundary DetectionHyolim Kang, Jinwoo Kim, Taehyun Kim, Seon Joo KimCVPR 2022 · 26 citations
- Scene Consistency Representation Learning for Video Scene SegmentationHaoqian Wu, Keyu Chen, Yanan Luo, Ruizhi Qiao et al.CVPR 2022 · 19 citations
Related papers
- Towards Global Video Scene Segmentation with Context-Aware TransformerYang Yang, Yurui Huang, Weili Guo, Baohua Xu et al.AAAI 2023 · 34 citations
- OS-MSL: One Stage Multimodal Sequential Link Framework for Scene Segmentation and ClassificationYe Liu, Lingfeng Qiao, Di Yin, Zhuoxuan Jiang et al.ACM MM 2022 · 5 citations
- Sharper and Faster mean Better: Towards More Efficient Vision-Language Model for Hour-scale Long Video UnderstandingDaoze Zhang, Yuze Zhao, Jintao Huang, Yingda ChenACL 2025 · 3 citations
- Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language ModelsNimrod Berman, Adam Botach, Emanuel Ben-Baruch, Shunit Haviv Hakimi et al.CVPR 2026 · 2 citations
- EDTER: Edge Detection with TransformerMengyang Pu, Yaping Huang, Yuming Liu, Qingji Guan et al.CVPR 2022 · 224 citations
