Neighbor Relations Matter in Video Scene Detection
Jiawei Tan, Hongxing Wang, Jiaxin Li, Zhilong Ou, Zhangbin Qian
Abstract
Video scene detection aims to temporally link shots for obtaining semantically compact scenes. It is essential for this task to capture scene-distinguishable affinity among shots by similarity assessment. However, most methods relies on ordinary shot-to-shot similarities, which may inveigle similar shots into being linked even though they are from different scenes, and meanwhile hinder dissimilar shots from being blended into a complete scene. In this paper, we propose NeighborNet to inject shot contexts into shot-to-shot similarities through carefully exploring the relations between semantic/temporal neighbors of shots over a local time period. In this way, shot-to-shot similarities are remeasured as semantic/temporal neighbor-aware similarities so that NeighborNet can learn context embedding into shot features using graph convolutional network. As a result, not only do the learned shot features suppress the affinity among similar shots from different scenes, but they also promote the affinity among dissimilar shots in the same scene. Experimental results on public benchmark datasets show that our proposed NeighborNet yields substantial improvements in video scene detection, especially outperforms released state-of-the-arts by at least 6% in Average Precision (AP). The code is available at https: //github.com/ExMorgan-Alter/NeighborNet .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 88e2a08e-226a-4872-a8ee-67ba50e84635Cited by top-tier papers2
- Video Color Grading via Look-Up Table GenerationSeunghyun Shin, Dongmin Shin, Jisu Shin, Hae-Gon Jeon et al.ICCV 2025 · 2 citations
- Video Scene Segmentation with Genre and Duration SignalsJungu Cho, Seong Jong Ha, Hae-Gon JeonICLR 2026
Builds on18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- How Attentive are Graph Attention Networks?Shaked Brody, Uri Alon, Eran YahavICLR 2022 · 1,717 citations
- Long Short-Term Transformer for Online Action DetectionMingze Xu, Yuanjun Xiong, Hao Chen, Xinyu Li et al.NeurIPS 2021 · 196 citations
Related papers
- Modality-Aware Shot Relating and Comparing for Video Scene DetectionJiawei Tan, Hongxing Wang, Kang Dang, Jiaxin Li et al.AAAI 2025 · 1 citation
- G-TAD: Sub-Graph Localization for Temporal Action DetectionMengmeng Xu, Chen Zhao, David S. Rojas, Ali K. Thabet et al.CVPR 2020
- Beyond Short-Term Snippet: Video Relation Detection With Spatio-Temporal Global ContextChenchen Liu, Yang Jin, Kehan Xu, Guoqiang Gong et al.CVPR 2020
- OS-MSL: One Stage Multimodal Sequential Link Framework for Scene Segmentation and ClassificationYe Liu, Lingfeng Qiao, Di Yin, Zhuoxuan Jiang et al.ACM MM 2022 · 5 citations
- Towards Global Video Scene Segmentation with Context-Aware TransformerYang Yang, Yurui Huang, Weili Guo, Baohua Xu et al.AAAI 2023 · 34 citations
