Shot Contrastive Self-Supervised Learning for Scene Boundary Detection
Shixing Chen, Xiaohan Nie, David Fan, Dongqing Zhang, Vimal Bhat, Raffay Hamid
摘要
Approach Overview -Representative frames of 10 shots from 2 different scenes of the movie Stuart Little are shown. The story-arch of each scene is distinguishable and semantically coherent. We consider similar nearby shots (e.g. 5 and 3) as augmented versions of each other. This augmentation approach is able to capitalize on the underlying film-production process and can encode the scenestructure better than the existing augmentation methods. Given a current shot (query) we find a similar shot (key) within its neighborhood and: (a) maximize the similarity between the query and the key, and (b) minimize the similarity of the query with randomly selected shots.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Towards Global Video Scene Segmentation with Context-Aware TransformerYang Yang, Yurui Huang, Weili Guo, Baohua Xu 等AAAI 2023 · 被引用 34 次
- UBoCo: Unsupervised Boundary Contrastive Learning for Generic Event Boundary DetectionHyolim Kang, Jinwoo Kim, Taehyun Kim, Seon Joo KimCVPR 2022 · 被引用 26 次
- Scene Consistency Representation Learning for Video Scene SegmentationHaoqian Wu, Keyu Chen, Yanan Luo, Ruizhi Qiao 等CVPR 2022 · 被引用 19 次
- Multimodal High-order Relation Transformer for Scene Boundary DetectionXi Wei, Zhangxiang Shi, Tianzhu Zhang, Xiaoyuan Yu 等ICCV 2023 · 被引用 7 次
- VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory BridgesYuxuan Wang, Yiqi Song, Cihang Xie, Yang Liu 等ICCV 2025 · 被引用 7 次
它引用的顶会 Paper8
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Data-Efficient Image Recognition with Contrastive Predictive CodingOlivier J. HénaffICML 2020 · 被引用 1,553 次
- Local Aggregation for Unsupervised Learning of Visual EmbeddingsChengxu Zhuang, Alex Lin Zhai, Daniel YaminsICCV 2019 · 被引用 462 次
- Self-supervised Co-Training for Video Representation LearningTengda Han, Weidi Xie, Andrew ZissermanNeurIPS 2020 · 被引用 405 次
- Enhancing Unsupervised Video Representation Learning by Decoupling the Scene and the MotionJinpeng Wang, Yuting Gao, Ke Li, Jianguo Hu 等AAAI 2021 · 被引用 70 次
相关 Paper
- Movies2Scenes: Using Movie Metadata to Learn Scene RepresentationShixing Chen, Chun-Hao Liu, Xiang Hao, Xiaohan Nie 等CVPR 2023
- STAGE: Storyboard-Anchored Generation for Cinematic Multi-shot NarrativePeixuan Zhang, Zijian Jia, Kaiqi Liu, Shuchen Weng 等CVPR 2026 · 被引用 25 次
- Similar Scenes Arouse Similar Emotions: Parallel Data Augmentation for Stylized Image CaptioningGuodun Li, Yuchen Zhai, Zehao Lin, Yin ZhangACM MM 2021 · 被引用 23 次
- OS-MSL: One Stage Multimodal Sequential Link Framework for Scene Segmentation and ClassificationYe Liu, Lingfeng Qiao, Di Yin, Zhuoxuan Jiang 等ACM MM 2022 · 被引用 5 次
- Depth-Guided Sparse Structure-from-Motion for Movies and TV ShowsSheng Liu, Xiaohan Nie, Raffay HamidCVPR 2022 · 被引用 12 次
