Towards Global Video Scene Segmentation with Context-Aware Transformer
Yang Yang, Yurui Huang, Weili Guo, Baohua Xu, Dingyin Xia
摘要
Videos such as movies or TV episodes usually need to divide the long storyline into cohesive units, i.e., scenes, to facilitate the understanding of video semantics. The key challenge lies in finding the boundaries of scenes by comprehensively considering the complex temporal structure and semantic information. To this end, we introduce a novel Context-Aware Transformer (CAT) with a self-supervised learning framework to learn high-quality shot representations, for generating well-bounded scenes. More specifically, we design the CAT with local-global self-attentions, which can effectively consider both the long-term and short-term context to improve the shot encoding. For training the CAT, we adopt the self-supervised learning schema. Firstly, we leverage shot-to-scene level pretext tasks to facilitate the pre-training with pseudo boundary, which guides CAT to learn the discriminative shot representations that maximize intra-scene similarity and inter-scene discrimination in an unsupervised manner. Then, we transfer contextual representations for fine-tuning the CAT with supervised data, which encourages CAT to accurately detect the boundary for scene segmentation. As a result, CAT is able to learn the context-aware shot representations and provides global guidance for scene segmentation. Our empirical analyses show that CAT can achieve state-of-the-art performance when conducting the scene segmentation task on the MovieNet dataset, e.g., offering 2.15 improvements on AP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- ShotBench: Expert-Level Cinematic Understanding in Vision-Language ModelsHongbo Liu, Jingwen He, Yi Jin, Dian Zheng 等NeurIPS 2025 · 被引用 24 次
- Modality-Aware Shot Relating and Comparing for Video Scene DetectionJiawei Tan, Hongxing Wang, Kang Dang, Jiaxin Li 等AAAI 2025 · 被引用 1 次
- Neighbor Relations Matter in Video Scene DetectionJiawei Tan, Hongxing Wang, Jiaxin Li, Zhilong Ou 等CVPR 2024 · 被引用 1 次
- Video Scene Segmentation with Genre and Duration SignalsJungu Cho, Seong Jong Ha, Hae-Gon JeonICLR 2026
- Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMsLucas Ventura, Antoine Yang, Cordelia Schmid, Gül VarolCVPR 2025
它引用的顶会 Paper8
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding 等ICCV 2019 · 被引用 709 次
- HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-trainingLinjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan 等EMNLP 2020 · 被引用 387 次
- Scene Consistency Representation Learning for Video Scene SegmentationHaoqian Wu, Keyu Chen, Yanan Luo, Ruizhi Qiao 等CVPR 2022 · 被引用 19 次
- DOMFN: A Divergence-Orientated Multi-Modal Fusion Network for Resume AssessmentYang Yang, Jingshuai Zhang, Fan Gao, Xiaoru Gao 等ACM MM 2022 · 被引用 12 次
相关 Paper
- Multimodal High-order Relation Transformer for Scene Boundary DetectionXi Wei, Zhangxiang Shi, Tianzhu Zhang, Xiaoyuan Yu 等ICCV 2023 · 被引用 7 次
- A Local-to-Global Approach to Multi-Modal Movie Scene SegmentationAnyi Rao, Linning Xu, Yu Xiong, Guodong Xu 等CVPR 2020
- Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language ModelsNimrod Berman, Adam Botach, Emanuel Ben-Baruch, Shunit Haviv Hakimi 等CVPR 2026 · 被引用 2 次
- No More Shortcuts: Realizing the Potential of Temporal Self-SupervisionIshan Rajendrakumar Dave, Simon Jenni, Mubarak ShahAAAI 2024 · 被引用 14 次
- Siamese Network with Interactive Transformer for Video Object SegmentationMeng Lan, Jing Zhang, Fengxiang He, Lefei ZhangAAAI 2022 · 被引用 41 次
