Multimodal High-order Relation Transformer for Scene Boundary Detection
Xi Wei, Zhangxiang Shi, Tianzhu Zhang, Xiaoyuan Yu, Lei Xiao
摘要
Scene boundary detection breaks down long videos into meaningful story-telling units and plays a crucial role in high-level video understanding. Despite significant advancements in this area, this task remains a challenging problem as it requires a comprehensive understanding of multimodal cues and high-level semantics. To tackle this issue, we propose a multimodal high-order relation transformer, which integrates a high-order encoder and an adaptive decoder in a unified framework. By modeling the mul-timodal cues and exploring similarities between the shots, the encoder is capable of capturing high-order relations between shots and extracting shot features with context semantics. By clustering the shots adaptively, the decoder can discover more universal switch pattern between successive scenes, thus helping scene boundary detection. Extensive experimental results on three standard benchmarks demonstrate that the proposed model performs favorably against state-of-the-art video scene detection methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Modality-Aware Shot Relating and Comparing for Video Scene DetectionJiawei Tan, Hongxing Wang, Kang Dang, Jiaxin Li 等AAAI 2025 · 被引用 1 次
- Neighbor Relations Matter in Video Scene DetectionJiawei Tan, Hongxing Wang, Jiaxin Li, Zhilong Ou 等CVPR 2024 · 被引用 1 次
- Video Scene Segmentation with Genre and Duration SignalsJungu Cho, Seong Jong Ha, Hae-Gon JeonICLR 2026
- Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMsLucas Ventura, Antoine Yang, Cordelia Schmid, Gül VarolCVPR 2025
它引用的顶会 Paper7
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Fast and Unsupervised Action Boundary Detection for Action SegmentationZexing Du, Xue Wang, Guoqing Zhou, Qing WangCVPR 2022 · 被引用 39 次
- UBoCo: Unsupervised Boundary Contrastive Learning for Generic Event Boundary DetectionHyolim Kang, Jinwoo Kim, Taehyun Kim, Seon Joo KimCVPR 2022 · 被引用 26 次
- Scene Consistency Representation Learning for Video Scene SegmentationHaoqian Wu, Keyu Chen, Yanan Luo, Ruizhi Qiao 等CVPR 2022 · 被引用 19 次
相关 Paper
- Towards Global Video Scene Segmentation with Context-Aware TransformerYang Yang, Yurui Huang, Weili Guo, Baohua Xu 等AAAI 2023 · 被引用 34 次
- OS-MSL: One Stage Multimodal Sequential Link Framework for Scene Segmentation and ClassificationYe Liu, Lingfeng Qiao, Di Yin, Zhuoxuan Jiang 等ACM MM 2022 · 被引用 5 次
- Sharper and Faster mean Better: Towards More Efficient Vision-Language Model for Hour-scale Long Video UnderstandingDaoze Zhang, Yuze Zhao, Jintao Huang, Yingda ChenACL 2025 · 被引用 3 次
- Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language ModelsNimrod Berman, Adam Botach, Emanuel Ben-Baruch, Shunit Haviv Hakimi 等CVPR 2026 · 被引用 2 次
- EDTER: Edge Detection with TransformerMengyang Pu, Yaping Huang, Yuming Liu, Qingji Guan 等CVPR 2022 · 被引用 224 次
