Generic Event Boundary Detection: A Benchmark for Event Segmentation
Mike Zheng Shou, Stan Weixian Lei, Weiyao Wang, Deepti Ghadiyaram, Matt Feiszli
Abstract
This paper presents a novel task together with a new benchmark for detecting generic, taxonomy-free event boundaries that segment a whole video into chunks. Conventional work in temporal video segmentation and action detection focuses on localizing pre-defined action categories and thus does not scale to generic videos. Cognitive Science has known since last century that humans consistently segment videos into meaningful temporal chunks. This segmentation happens naturally, without pre-defined event categories and without being explicitly asked to do so. Here, we repeat these cognitive experiments on mainstream CV datasets; with our novel annotation guideline which addresses the complexities of taxonomy-free event boundary annotation, we introduce the task of Generic Event Boundary Detection (GEBD) and the new benchmark Kinetics-GEBD. We view GEBD as an important stepping stone towards understanding the video as a whole, and believe it has been previously neglected due to a lack of proper task definition and annotations. Through experiment and human study we demonstrate the value of the annotations. Further, we benchmark supervised and un-supervised GEBD approaches on the TAPOS dataset and our Kinetics-GEBD. We release our annotations and baseline codes at CVPR'21 LOVEU Challenge: https://sites.google.com/ view/loveucvpr21 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3d45f0fe-7edf-4cfa-8b34-a4e0989ae6aaCited by top-tier papers21
- Anticipative Video TransformerRohit Girdhar, Kristen GraumanICCV 2021 · 270 citations
- Knowing Where to Focus: Event-aware Transformer for Video GroundingJinhyun Jang, Jungin Park, Jin Kim, Hyeongjun Kwon et al.ICCV 2023 · 103 citations
- MovieChat: From Dense Token to Sparse Memory for Long Video UnderstandingEnxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang et al.CVPR 2024 · 95 citations
- Equivariant Similarity for Vision-Language Foundation ModelsTan Wang, Kevin Lin, Linjie Li, Chung-Ching Lin et al.ICCV 2023 · 67 citations
- Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal UnderstandingYunlong Tang, Daiki Shimada, Jing Bi, Mingqian Feng et al.AAAI 2025 · 29 citations
Builds on8
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding et al.ICCV 2019 · 709 citations
- HACS: Human Action Clips and Segments Dataset for Recognition and Temporal LocalizationHang Zhao, Antonio Torralba, Lorenzo Torresani, Zhicheng YanICCV 2019 · 298 citations
- Weakly Supervised Energy-Based Learning for Action SegmentationJun Li, Peng Lei, Sinisa TodorovicICCV 2019 · 109 citations
Related papers
- Generic Event Boundary Detection via Denoising DiffusionJaejun Hwang, Dayoung Gong, Manjin Kim, Minsu ChoICCV 2025
- Progressive Attention on Multi-Level Dense Difference Maps for Generic Event Boundary DetectionJiaqi Tang, Zhaoyang Liu, Chen Qian, Wayne Wu et al.CVPR 2022 · 15 citations
- Online Generic Event Boundary DetectionHyungrok Jung, Daneul Kim, Seunggyun Lim, Jeany Son et al.ICCV 2025
- End-to-End Compressed Video Representation Learning for Generic Event Boundary DetectionCongcong Li, Xinyao Wang, Longyin Wen, Dexiang Hong et al.CVPR 2022 · 18 citations
- Rethinking the Architecture Design for Efficient Generic Event Boundary DetectionZiwei Zheng, Zechuan Zhang, Yulin Wang, Shiji Song et al.ACM MM 2024 · 1 citation
